TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 30 – Oct 6, 2024
本周最热99

Emu3: Next-Token Prediction is All You Need

Xinlong Wang, Xiaosong Zhang, Zhengxiong Luo +22 authors

Emu3, a transformer-based multimodal model trained exclusively with next-token prediction, outperforms existing diffusion and compositional models in generation and perception tasks.

next-token predictionmultimodal modelstransformerdiscrete spaceHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

MIO: A Foundation Model on Multimodal Tokens

Zekun Wang, King Zhu, Chunpu Xu +14 authors

MIO, a novel foundation model, achieves competitive performance in multimodal tasks through an end-to-end, autoregressive approach using causal multimodal modeling.

53MIOmultimodal tokensHF ↗arXiv ↗
07

Depth Pro: Sharp Monocular Metric Depth in Less Than a Second

Aleksei Bochkovskii, Amaël Delaunoy, Hugo Germain +4 authors

Depth Pro is a fast and accurate zero-shot monocular depth estimation model that generates high-resolution depth maps using a multi-scale transformer and a combined real-synthetic training protocol.

43zero-shotmonocular depth estimationHF ↗arXiv ↗
08

Video Instruction Tuning With Synthetic Data

Yuanhan Zhang, Jinming Wu, Wei Li +4 authors

A synthetic dataset, LLaVA-Video-178K, and a new video large multimodal model, LLaVA-Video, achieve strong performance in video instruction-following tasks.

41video large multimodal modelssynthetic datasetHF ↗arXiv ↗
11

LLaVA-Critic: Learning to Evaluate Multimodal Models

Tianyi Xiong, Xiyao Wang, Dong Guo +5 authors

LLaVA-Critic, an open-source large multimodal model, effectively evaluates multimodal tasks and provides reliable scores, surpassing GPT models, and enhances preference learning for model alignment.

37large multimodal modelLMMHF ↗arXiv ↗
12

Contrastive Localized Language-Image Pre-Training

Hong-You Chen, Zhengfeng Lai, Haotian Zhang +7 authors

CLOC enhances CLIP's localization capabilities by introducing region-text contrastive loss and promptable embeddings, improving regional image representation for multimodal large language models.

36Contrastive Language-Image Pre-training (CLIP)multimodal large language models (MLLMs)HF ↗arXiv ↗
16

PHI-S: Distribution Balancing for Label-Free Multi-Teacher Distillation

Mike Ranzinger, Jon Barker, Greg Heinrich +3 authors

Heterogeneous multi-teacher knowledge distillation improves visual foundation models by optimizing activation statistics and alignment, with PHI Standardization shown to achieve the best student model quality.

33heterogeneous multi-teacher knowledge distillationactivation statisticsHF ↗arXiv ↗
18

Large Language Models as Markov Chains

Oussama Zekri, Ambroise Odonnat, Abdelhakim Benechehab +3 authors

Theoretical analysis of large language models explores their inference capabilities and generalization through equivalence with Markov chains, providing pre-training and in-context bounds.

32autoregressive language modelsMarkov chainsHF ↗arXiv ↗
20

A Survey on the Honesty of Large Language Models

Siheng Li, Cheng Yang, Taiqiang Wu +12 authors

Honesty is a fundamental principle for aligning large language models (LLMs) with human values, requiring these models to recognize what they know and don't know and be able to faithfully express their knowledge. Despite promising, current LLMs still exhibit significant dishonest behaviors, such as confidently presenting wrong answers or failing to express what they know. In addition, research on the honesty of LLMs also faces challenges, including varying definitions of honesty, difficulties in distinguishing between known and unknown knowledge, and a lack of comprehensive understanding of related research. To address these issues, we provide a survey on the honesty of LLMs, covering its clarification, evaluation approaches, and strategies for improvement. Moreover, we offer insights for future research, aiming to inspire further exploration in this important area.

31HF ↗arXiv ↗
21

Not All LLM Reasoners Are Created Equal

Arian Hosseini, Alessandro Sordoni, Daniel Toyama +2 authors

LLMs display a significant reasoning gap when solving dependent math problems, which is more pronounced in smaller and math-specialized models, and is attributed to distraction from additional context and poor second-hop reasoning.

28HF ↗arXiv ↗
22

LEOPARD : A Vision Language Model For Text-Rich Multi-Image Tasks

Mengzhao Jia, Wenhao Yu, Kaixin Ma +6 authors

A model designed for multimodal tasks involving multiple text-rich images addresses dataset scarcity and visual feature balancing, outperforming existing models in specialized and general evaluations.

28multimodal large language modelshigh-quality multimodal instruction-tuning dataHF ↗arXiv ↗
23

Hyper-Connections

Defa Zhu, Hongzhi Huang, Zihao Huang +5 authors

Hyper-connections, an alternative to residual connections, improve performance in pre-training large language models and vision tasks by dynamically adjusting feature connections.

28hyper-connectionsresidual connectionsHF ↗arXiv ↗
30

Contextual Document Embeddings

John X. Morris, Alexander M. Rush

Contextualized document embeddings that consider neighboring documents improve retrieval performance over traditional biencoders, achieving state-of-the-art results on the MTEB benchmark.

23dense document embeddingsneural retrievalHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号