TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

121

Gemma 2: Improving Open Language Models at a Practical Size

Gemma Team, Morgane Riviere, Shreya Pathak +193 authors

Gemma 2 introduces improvements in the Transformer architecture through interleaving local-global attentions and group-query attention, showcasing superior performance relative to its size.

79Transformer architectureinterleaving local-global attentionsHF ↗arXiv ↗
124

OmniFusion Technical Report

Elizaveta Goncharova, Anton Razzhigaev, Matvey Mikhalchuk +6 authors

The OmniFusion model, leveraging pretrained LLMs and visual adapters, demonstrates superior performance across multiple visual-language benchmarks compared to open-source alternatives.

78multimodal architecturesOmniFusion modelHF ↗arXiv ↗
126

LLaMA-Mesh: Unifying 3D Mesh Generation with Language Models

Zhengyi Wang, Jonathan Lorraine, Yikai Wang +4 authors

The work demonstrates the capability of LLMs to generate 3D meshes from text by introducing a novel approach to tokenize 3D mesh data, allowing the unification of 3D and text modalities without expanding the model's vocabulary.

78large language modelsLLMsHF ↗arXiv ↗
127

Generative World Explorer

Taiming Lu, Tianmin Shu, Alan Yuille +2 authors

Generative World Explorer (Genex) enables agents to mentally explore 3D environments using imagined observations, updating their beliefs without physical exploration to improve decision-making.

77Generative World ExplorerGenexHF ↗arXiv ↗
131

NVLM: Open Frontier-Class Multimodal LLMs

Wenliang Dai, Nayeon Lee, Boxin Wang +7 authors

NVML 1.0, a family of multimodal large language models, achieves state-of-the-art results in vision-language tasks by combining text-only and multimodal training, utilizing a new architecture and dataset strategy.

75multimodal large language modelsdecoder-only multimodal LLMsHF ↗arXiv ↗
135

STIV: Scalable Text and Image Conditioned Video Generation

Zongyu Lin, Wei Liu, Chen Chen +14 authors

STIV, a text-image-conditioned video generation method integrating Diffusion Transformer and classifier-free guidance, achieves state-of-the-art performance in text-to-video, text-image-to-video, and image-to-video tasks.

74video generationmodel architecturesHF ↗arXiv ↗
137

LLM Agent Operating System

Kai Mei, Zelong Li, Shuyuan Xu +3 authors

AIOS, an operating system embedding large language models, addresses resource allocation, context switching, and concurrency challenges for intelligent agents, demonstrating reliability and efficiency.

73large language modelintelligent agentsHF ↗arXiv ↗
139

PaliGemma: A versatile 3B VLM for transfer

Lucas Beyer, Andreas Steiner, André Susano Pinto +32 authors

PaliGemma, a versatile Vision-Language Model based on SigLIP-So400m and Gemma-2B, demonstrates strong performance across numerous open-world tasks, including specialized areas like remote sensing and segmentation.

73Vision-Language ModelSigLIP-So400mHF ↗arXiv ↗
142

RAFT: Adapting Language Model to Domain Specific RAG

Tianjun Zhang, Shishir G. Patil, Naman Jain +4 authors

Retrieval Augmented FineTuning (RAFT) enhances pre-trained large language models' in-domain performance by filtering out irrelevant documents and citing relevant passages.

72Retrieval Augmented FineTuningRAFTHF ↗arXiv ↗
143

Progressive Multimodal Reasoning via Active Retrieval

Guanting Dong, Chenghao Zhang, Mengjie Deng +3 authors

AR-MCTS enhances multimodal large language models' reasoning capabilities through active retrieval, Monte Carlo Tree Search, and a process reward model, improving performance across multimodal reasoning tasks.

72Active RetrievalMonte Carlo Tree SearchHF ↗arXiv ↗
146

Genie: Generative Interactive Environments

Jake Bruce, Michael Dennis, Ashley Edwards +22 authors

Genie, a 11B parameter unsupervised generative model, creates action-controllable virtual worlds from unlabelled videos using spatiotemporal tokenization and autoregressive dynamics, enabling agent training from unseen video behaviors.

72spatiotemporal video tokenizerautoregressive dynamics modelHF ↗arXiv ↗
147

Towards a Unified View of Preference Learning for Large Language Models: A Survey

Bofei Gao, Feifan Song, Yibo Miao +21 authors

Large Language Models (LLMs) exhibit remarkably powerful capabilities. One of the crucial factors to achieve success is aligning the LLM's output with human preferences. This alignment process often requires only a small amount of data to efficiently enhance the LLM's performance. While effective, research in this area spans multiple domains, and the methods involved are relatively complex to understand. The relationships between different methods have been under-explored, limiting the development of the preference alignment. In light of this, we break down the existing popular alignment strategies into different components and provide a unified framework to study the current alignment strategies, thereby establishing connections among them. In this survey, we decompose all the strategies in preference learning into four components: model, data, feedback, and algorithm. This unified view offers an in-depth understanding of existing alignment algorithms and also opens up possibilities to synergize the strengths of different strategies. Furthermore, we present detailed working examples of prevalent existing algorithms to facilitate a comprehensive understanding for the readers. Finally, based on our unified perspective, we explore the challenges and future research directions for aligning large language models with human preferences.

72HF ↗arXiv ↗
150

Kvasir-VQA: A Text-Image Pair GI Tract Dataset

Sushant Gautam, Andrea Storås, Cise Midoglu +4 authors

Kvasir-VQA is a dataset with question-and-answer annotations for GI diagnostics, supporting image captioning, VQA, synthetic image generation, object detection, and classification.

71Visual Question Answering (VQA)image captioningHF ↗arXiv ↗
5 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号