TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

May 13 – May 19, 2024
本周最热134

Chameleon: Mixed-Modal Early-Fusion Foundation Models

Chameleon Team

Chameleon is a mixed-modal early-fusion token-based model that achieves state-of-the-art performance across various tasks, including image captioning, text generation, and long-form mixed-modal generation, using a unified architecture.

early-fusiontoken-basedmixed-modal modelsvisual question answeringHF ↗arXiv ↗

29 篇论文 · 按点赞排序

02

LoRA Learns Less and Forgets Less

Dan Biderman, Jose Gonzalez Ortiz, Jacob Portes +9 authors

LoRA, a parameter-efficient finetuning method for large language models, underperforms full finetuning in target domains but provides better regularization and maintains diverse generation compared to other techniques.

91Low-Rank AdaptationLoRAHF ↗arXiv ↗
03

What matters when building vision-language models?

Hugo Laurençon, Léo Tronchon, Matthieu Cord +1 authors

Idefics2, a vision-language model with 8 billion parameters, achieves state-of-the-art performance on multimodal benchmarks through extensive experimental validation.

77vision-language modelslarge language modelsHF ↗arXiv ↗
04

RLHF Workflow: From Reward Modeling to Online RLHF

Hanze Dong, Wei Xiong, Bo Pang +7 authors

Online iterative reinforcement learning from human feedback achieves state-of-the-art performance in large language models using open-source datasets and proxy preference models.

71RLHFOnline Iterative RLHFHF ↗arXiv ↗
06

SUTRA: Scalable Multilingual Language Model Architecture

Abhijit Bendale, Michael Sapienza, Steven Ripplinger +3 authors

SUTRA, a multilingual Large Language Model architecture, achieves superior performance on multilingual tasks by decoupling conceptual understanding from language-specific processing using a Mixture of Experts framework.

37Multilingual Large Language ModelMixture of Experts frameworkHF ↗arXiv ↗
09

Many-Shot In-Context Learning in Multimodal Foundation Models

Yixing Jiang, Jeremy Irvin, Ji Hun Wang +3 authors

Multimodal foundation models exhibit significant performance improvements with many-shot in-context learning compared to few-shot, with Gemini 1.5 Pro showing higher data efficiency and better batch processing capabilities across various domains.

29few-shot in-context learningGPT-4oHF ↗arXiv ↗
13

Toon3D: Seeing Cartoons from a New Perspective

Ethan Weber, Riley Peterlinz, Rohan Mathur +3 authors

A pipeline combining camera pose estimation and image deformation corrects 2D drawing inconsistencies in hand-drawn cartoons to recover plausible 3D structures for novel-view synthesis.

20camera pose estimationimage deformationHF ↗arXiv ↗

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号