TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

181

OtterHD: A High-Resolution Multi-modality Model

Bo Li, Peiyuan Zhang, Jingkang Yang +3 authors

OtterHD-8B, an advanced multimodal model from Fuyu-8B, excels in processing high-resolution inputs and discerning detailed spatial relationships through MagnifierBench, an evaluation framework highlighting the importance of vision encoder flexibility.

34multimodal modelhigh-resolution visual inputsHF ↗arXiv ↗
183

LARP: Language-Agent Role Play for Open-World Games

Ming Yan, Ruihao Li, Hao Zhang +3 authors

LARP is a language agent framework for open-world games, incorporating memory and decision-making capabilities to enhance interaction and coherence in a diverse range of applications.

34cognitive architecturememory processingHF ↗arXiv ↗
186

Shepherd: A Critic for Language Model Generation

Tianlu Wang, Ping Yu, Xiaoqing Ellen Tan +7 authors

Shepherd, a small language model tuned for critique, outperforms or ties with larger models like ChatGPT in refining language model outputs using a high-quality feedback dataset.

33language modelcritiqueHF ↗arXiv ↗
187

OctoPack: Instruction Tuning Code Large Language Models

Niklas Muennighoff, Qian Liu, Armel Zebaze +7 authors

Instruction tuning using Git commits improves performance on natural language and coding tasks compared to other benchmarks, with models achieving state-of-the-art results on expanded HumanEvalPack.

33instruction tuningcodeHF ↗arXiv ↗
189

Text-to-3D using Gaussian Splatting

Zilong Chen, Feng Wang, Huaping Liu

Gaussian Splatting based text-to-3D generation improves geometry accuracy and fidelity through progressive optimization, including geometry optimization and appearance refinement.

33Gaussian Splatting3D priorHF ↗arXiv ↗
193

FP8-LM: Training FP8 Large Language Models

Houwen Peng, Kan Wu, Yixuan Wei +17 authors

A new FP8 automatic mixed-precision framework for training large language models reduces memory usage and increases speed compared to BF16 and Nvidia Transformer Engine.

33FP8low-bit data formatsHF ↗arXiv ↗
196

Analyzing and Improving the Training Dynamics of Diffusion Models

Tero Karras, Miika Aittala, Jaakko Lehtinen +3 authors

Modifications to network layers in the ADM diffusion model architecture improve training stability and synthesis quality, reducing FID from 2.41 to 1.81, and a novel method for post-hoc EMA parameter tuning is introduced.

33diffusion modelsADM diffusion modelHF ↗arXiv ↗
201

Aligning Large Multimodal Models with Factually Augmented RLHF

Zhiqing Sun, Sheng Shen, Shengcao Cao +9 authors

The paper presents Factually Augmented RLHF to address multimodal misalignment in large multimodal models, significantly improving hallucination reduction and overall performance on vision-language tasks compared to existing methods.

32Reinforcement Learning from Human Feedback (RLHF)vision-language alignmentHF ↗arXiv ↗
204

CogAgent: A Visual Language Model for GUI Agents

Wenyi Hong, Weihan Wang, Qingsong Lv +8 authors

CogAgent, a visual language model with strong GUI understanding and navigation capabilities, outperforms LLM-based methods in both PC and Android GUI tasks using only screenshots.

32visual language modelGUI understandingHF ↗arXiv ↗
205

FaceStudio: Put Your Face Everywhere in Seconds

Yuxuan Yan, Chi Zhang, Rui Wang +3 authors

A hybrid guidance framework for identity-preserving image synthesis efficiently generates stylistic portraits by combining stylized images, facial images, and textual prompts.

32identity-preserving synthesisTextual InversionHF ↗arXiv ↗
206

Relightable Gaussian Codec Avatars

Shunsuke Saito, Gabriel Schwartz, Tomas Simon +2 authors

Relightable Gaussian Codec Avatars model high-fidelity head avatars with real-time relighting capabilities using 3D Gaussians for geometry and learnable radiance transfer for appearance.

323D Gaussiansrelightable appearance modelHF ↗arXiv ↗
209

Effective Long-Context Scaling of Foundation Models

Wenhan Xiong, Jingyu Liu, Igor Molybog +18 authors

Long-context LLMs achieve significant advancements in handling extended contexts through continual pretraining and efficient instruction tuning, surpassing existing models on various benchmarks.

31LLMslong-contextHF ↗arXiv ↗
7 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号