TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

598 篇论文 · 按点赞排序

571

Can large language models explore in-context?

Akshay Krishnamurthy, Keegan Harris, Dylan J. Foster +2 authors

Experimentation with LLMs in multi-armed bandit environments reveals that robust exploration requires specific prompts, external summarization of history, or algorithmic interventions.

33Large Language ModelsLLMsHF ↗arXiv ↗
572

Larimar: Large Language Models with Episodic Memory Control

Payel Das, Subhajit Chaudhury, Elliot Nelson +10 authors

Larimar, a brain-inspired architecture with distributed episodic memory, enhances LLMs for efficient, fast, and accurate knowledge updates, fact editing, and context length generalization.

33Larimarbrain-inspired architectureHF ↗arXiv ↗
574

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

William Brandon, Mayank Mishra, Aniruddha Nrusimha +2 authors

Cross-Layer Attention modifies transformer-based autoregressive large language models to reduce KV cache size while maintaining accuracy, enabling longer sequence lengths and larger batch sizes during inference.

33Key-value cachingtransformer-based autoregressive large language modelsHF ↗arXiv ↗
579

Best Practices and Lessons Learned on Synthetic Data for Language Models

Ruibo Liu, Jerry Wei, Fangyu Liu +8 authors

The success of AI models relies on the availability of large, diverse, and high-quality datasets, which can be challenging to obtain due to data scarcity, privacy concerns, and high costs. Synthetic data has emerged as a promising solution by generating artificial data that mimics real-world patterns. This paper provides an overview of synthetic data research, discussing its applications, challenges, and future directions. We present empirical evidence from prior art to demonstrate its effectiveness and highlight the importance of ensuring its factuality, fidelity, and unbiasedness. We emphasize the need for responsible use of synthetic data to build more powerful, inclusive, and trustworthy language models.

32HF ↗arXiv ↗
581

FlashSpeech: Efficient Zero-Shot Speech Synthesis

Zhen Ye, Zeqian Ju, Haohe Liu +10 authors

FlashSpeech, using latent consistency model and adversarial consistency training, achieves fast and high-quality zero-shot speech synthesis with enhanced prosody generation.

32latent consistency modeladversarial consistency trainingHF ↗arXiv ↗
582

mPLUG-DocOwl 1.5: Unified Structure Learning for OCR-free Document Understanding

Anwen Hu, Haiyang Xu, Jiabo Ye +8 authors

A Unified Structure Learning framework is introduced to enhance Visual Document Understanding by integrating structure-aware parsing and multi-grained text localization tasks using a vision-to-text module H-Reducer, achieving state-of-the-art results on multiple benchmarks.

32Multimodal Large Language ModelsVisual Document UnderstandingHF ↗arXiv ↗
583

Many-Shot In-Context Learning in Multimodal Foundation Models

Yixing Jiang, Jeremy Irvin, Ji Hun Wang +3 authors

Multimodal foundation models exhibit significant performance improvements with many-shot in-context learning compared to few-shot, with Gemini 1.5 Pro showing higher data efficiency and better batch processing capabilities across various domains.

32few-shot in-context learningGPT-4oHF ↗arXiv ↗
588

NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Gerald Shen, Zhilin Wang, Olivier Delalleau +10 authors

NeMo-Aligner is a toolkit for aligning large language models with human values using scalable techniques like RLHF, DPO, SteerLM, and SPIN, optimized for use with hundreds of GPUs and parameter-efficient fine-tuning.

30NeMo-AlignerReinforcement Learning from Human Feedback (RLHF)HF ↗arXiv ↗
592

Self-Play Preference Optimization for Language Model Alignment

Yue Wu, Zhiqing Sun, Huizhuo Yuan +3 authors

A self-play method called SPPO for language model alignment achieves state-of-the-art performance by approximating Nash equilibrium policy in a constant-sum game setting, outperforming other approaches with limited data.

29reinforcement learning from human feedbackparametric modelsHF ↗arXiv ↗
593

Octo: An Open-Source Generalist Robot Policy

Octo Model Team, Dibya Ghosh, Homer Walke +15 authors

Octo, a large transformer-based policy trained on extensive robotic datasets, demonstrates versatility and efficient fine-tuning for diverse robotic platforms and tasks.

29transformer-based policyOpen X-Embodiment datasetHF ↗arXiv ↗
594

Imp: Highly Capable Large Multimodal Models for Mobile Devices

Zhenwei Shao, Zhou Yu, Jun Yu +5 authors

A systematic study of lightweight large multimodal models (LMMs) led to the development of the Imp family, which achieves superior performance compared to larger models and is capable of high-speed inference on mobile devices.

29large language modelslarge multimodal modelsHF ↗arXiv ↗
597

The Road Less Scheduled

Aaron Defazio, Xingyu, Yang +4 authors

A new schedule-free method achieves state-of-the-art performance in optimization without requiring a stopping time or additional hyper-parameters, by unifying scheduling and iterate averaging.

27learning rate schedulesoptimization stopping stepHF ↗arXiv ↗
20 / 20

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号