TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

403 篇论文 · 按点赞排序

337

RWKV: Reinventing RNNs for the Transformer Era

Bo Peng, Eric Alcaide, Quentin Anthony +27 authors

A new architecture, RWKV, combines the parallelizable training of Transformers with the efficient inference of RNNs, achieving linear scaling and matching performance.

20Transformersrecurrent neural networks (RNNs)HF ↗arXiv ↗
343

Conditional Diffusion Distillation

Kangfu Mei, Mauricio Delbracio, Hossein Talebi +3 authors

A novel single-stage distillation method for generative diffusion models reduces sampling time while maintaining performance across tasks like super-resolution and image editing.

19generative diffusion modelstext-to-image generationHF ↗arXiv ↗
344

Augmenting Language Models with Long-Term Memory

Weizhi Wang, Li Dong, Hao Cheng +4 authors

A framework called LongMem enables large language models to utilize long-term memory, overcoming input length limitations and improving performance on long-context tasks.

19language modelslong-term memoryHF ↗arXiv ↗
346

h2oGPT: Democratizing Large Language Models

Arno Candel, Jon McKinney, Philipp Singer +12 authors

h2oGPT provides open-source, fine-tuned LLMs based on Generative Pretrained Transformers with 100% private document search capabilities.

19Generative Pretrained TransformersLLMsHF ↗arXiv ↗
350

CausalLM is not optimal for in-context learning

Nan Ding, Tomer Levinboim, Jialin Wu +2 authors

Theoretical analysis shows that prefix language models outperform causal language models in in-context learning by converging to optimal solutions, while causal models exhibit dynamics similar to online gradient descent.

19transformerin-context learningHF ↗arXiv ↗
354

DreamHuman: Animatable 3D Avatars from Text

Nikos Kolotouros, Thiemo Alldieck, Andrei Zanfir +3 authors

DreamHuman generates realistic animatable 3D human avatars using text by integrating text-to-image synthesis, neural radiance fields, and statistical human body models.

17text-to-3D methodsneural radiance fieldsHF ↗arXiv ↗
355

HomeRobot: Open-Vocabulary Mobile Manipulation

Sriram Yenamandra, Arun Ramachandran, Karmesh Yadav +15 authors

The HomeRobot OVMM benchmark evaluates robots' ability to manipulate unseen objects in new environments using a combination of simulated and real-world testing.

17Open-Vocabulary Mobile ManipulationpercpetionHF ↗arXiv ↗
356

Long-range Language Modeling with Self-retrieval

Ohad Rubin, Jonathan Berant

The Retrieval-Pretrained Transformer jointly trains a retrieval-augmented language model to model long texts by fusing retrieved information for improved prediction.

17retrieval-augmented language modelsRetrieval-Pretrained TransformerHF ↗arXiv ↗
357

Scaling MLPs: A Tale of Inductive Bias

Gregor Bachmann, Sotiris Anagnostidis, Thomas Hofmann

MLPs achieve competitive performance on vision tasks with large-scale pre-training, challenging the narrative that inductive bias is necessary for high accuracy.

17multi-layer perceptron (MLP)inductive biasHF ↗arXiv ↗
358

Scalable 3D Captioning with Pretrained Models

Tiange Luo, Chris Rockwell, Honglak Lee +1 authors

Cap3D generates high-quality descriptive text for 3D objects using pretrained models and datasets, surpassing human performance in quality, cost, and speed.

17image captioningimage-text alignmentHF ↗arXiv ↗
360

SoundStorm: Efficient Parallel Audio Generation

Zalán Borsos, Matt Sharifi, Damien Vincent +3 authors

SoundStorm, a non-autoregressive audio generation model, delivers high-quality and consistent audio two orders of magnitude faster than autoregressive methods.

15SoundStormnon-autoregressiveHF ↗arXiv ↗
12 / 14

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号