TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Aug 26 – Sep 1, 2024
本周最热144

Writing in the Margins: Better Inference Pattern for Long Context Retrieval

Melisa Russak, Umar Jamil, Christopher Bryant +4 authors

Writing in the Margins (WiM) enhances Large Language Models' performance on long input sequences and retrieval tasks by using chunked prefill and marginal information, improving accuracy and F1-score without fine-tuning.

Large Language Modelschunked prefillkey-value cachesegment-wise inferenceHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Diffusion Models Are Real-Time Game Engines

Dani Valevski, Yaniv Leviathan, Moab Arar +1 authors

GameNGen, a neural model-powered game engine, simulates high-quality gameplay in real-time using a diffusion model conditioned on past frames and actions.

127neural modelreal-time interactionHF ↗arXiv ↗
04

Law of Vision Representation in MLLMs

Shijia Yang, Bohan Zhai, Quanzeng You +3 authors

Correlation between cross-modal alignment and vision representation improves performance in multimodal large language models, enabling identification and training of optimal vision representation with reduced computational cost.

95cross-modal alignmentvision representationHF ↗arXiv ↗
10

Foundation Models for Music: A Survey

Yinghao Ma, Anders Øland, Anton Ragni +40 authors

A review of foundation models in music, including large language models and latent diffusion models, highlights their impact, architectural choices, and the need for ethical considerations in music applications.

44large language modelslatent diffusion modelsHF ↗arXiv ↗
19

Text2SQL is Not Enough: Unifying AI and Databases with TAG

Asim Biswal, Liana Patel, Siddarth Jha +5 authors

A unified Table-Augmented Generation (TAG) paradigm addresses limitations of existing Text2SQL and retrieval-based methods, enabling a broader range of natural language queries over databases with improved accuracy and flexibility.

26Text2SQLRetrieval-Augmented GenerationHF ↗arXiv ↗
28

LLaVA-MoD: Making LLaVA Tiny via MoE Knowledge Distillation

Fangxun Shu, Yue Liao, Le Zhuo +13 authors

LLaVA-MoD employs a sparse Mixture of Experts and progressive knowledge transfer to efficiently train small-scale multimodal language models with limited computational resources while outperforming larger models.

21Mixture of Experts (MoE)mimic distillationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号