TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

92

Kandinsky 3.0 Technical Report

Vladimir Arkhipkin, Andrei Filatov, Viacheslav Vasilev +6 authors

Kandinsky 3.0, a large-scale text-to-image model based on latent diffusion, improves quality and realism through a larger architecture and advanced text understanding.

45latent diffusionU-NetHF ↗arXiv ↗
101

Judging LLM-as-a-judge with MT-Bench and Chatbot Arena

Lianmin Zheng, Wei-Lin Chiang, Ying Sheng +10 authors

Using strong large language models as judges for evaluating other LLM-based chat assistants achieves high agreement with human preferences, offering a scalable and explainable solution compared to traditional benchmarks.

43large language modelLLMHF ↗arXiv ↗
104

In-Context Learning Creates Task Vectors

Roee Hendel, Mor Geva, Amir Globerson

In-Context Learning in Large Language Models can be understood as compressing a training set into a task vector that modulates a transformer for output generation.

43in-context learninglarge language modelsHF ↗arXiv ↗
105

Agents: An Open-source Framework for Autonomous Language Agents

Wangchunshu Zhou, Yuchen Eleanor Jiang, Long Li +14 authors

Agents is an open-source library that facilitates the creation, customization, and deployment of autonomous language agents with features like planning, memory, and tool usage, aiming to make advances in large language models accessible to a broader audience.

43large language modelsautonomous language agentsHF ↗arXiv ↗
108

Self-Alignment with Instruction Backtranslation

Xian Li, Ping Yu, Chunting Zhou +5 authors

A scalable instruction-following language model is built using auto-labelling and iterative self-augmentation and self-curation, outperforming other LLaMa-based models on Alpaca.

43instruction backtranslationlanguage modelHF ↗arXiv ↗
109

Brain2Music: Reconstructing Music from Human Brain Activity

Timo I. Denk, Yu Takagi, Takuya Matsuyama +4 authors

A method reconstructs music from fMRI data using the MusicLM model, matching semantic properties of original stimuli and identifying brain regions involved in processing music descriptions.

42functional magnetic resonance imagingfMRIHF ↗arXiv ↗
116

3D-LLM: Injecting the 3D World into Large Language Models

Yining Hong, Haoyu Zhen, Peihao Chen +4 authors

A new family of 3D-LLMs is introduced to perform 3D-related tasks by leveraging 3D point clouds and features, outperforming state-of-the-art baselines in tasks such as 3D question answering and captioning.

40LLMsVision-Language ModelsHF ↗arXiv ↗
117

4K4D: Real-Time 4D View Synthesis at 4K Resolution

Zhen Xu, Sida Peng, Haotong Lin +5 authors

4K4D, a 4D point cloud representation with a hybrid appearance model and differentiable depth peeling algorithm, achieves high-speed and high-quality dynamic view synthesis at 4K resolution.

404D point cloud representationhardware rasterizationHF ↗arXiv ↗
118

Table-GPT: Table-tuned GPT for Diverse Table Tasks

Peng Li, Yeye He, Dror Yashar +6 authors

A new table-tuning paradigm improves language models' understanding and generalization in table-related tasks by fine-tuning them with synthesized table data, resulting in enhanced performance.

40table-tuningtable-understandingHF ↗arXiv ↗
4 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号