TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

50 篇论文 · 按点赞排序

35

Extending Llama-3's Context Ten-Fold Overnight

Peitian Zhang, Ninglu Shao, Zheng Liu +4 authors

Llama-3-8B-Instruct's context length is extended from 8K to 80K using QLoRA fine-tuning with minimal additional training samples, demonstrating significant potential for further context extension with increased computational resources.

34QLoRA fine-tuningHF ↗arXiv ↗
36

Reducing Transformer Key-Value Cache Size with Cross-Layer Attention

William Brandon, Mayank Mishra, Aniruddha Nrusimha +2 authors

Cross-Layer Attention modifies transformer-based autoregressive large language models to reduce KV cache size while maintaining accuracy, enabling longer sequence lengths and larger batch sizes during inference.

33Key-value cachingtransformer-based autoregressive large language modelsHF ↗arXiv ↗
38

Many-Shot In-Context Learning in Multimodal Foundation Models

Yixing Jiang, Jeremy Irvin, Ji Hun Wang +3 authors

Multimodal foundation models exhibit significant performance improvements with many-shot in-context learning compared to few-shot, with Gemini 1.5 Pro showing higher data efficiency and better batch processing capabilities across various domains.

32few-shot in-context learningGPT-4oHF ↗arXiv ↗
43

NeMo-Aligner: Scalable Toolkit for Efficient Model Alignment

Gerald Shen, Zhilin Wang, Olivier Delalleau +10 authors

NeMo-Aligner is a toolkit for aligning large language models with human values using scalable techniques like RLHF, DPO, SteerLM, and SPIN, optimized for use with hundreds of GPUs and parameter-efficient fine-tuning.

30NeMo-AlignerReinforcement Learning from Human Feedback (RLHF)HF ↗arXiv ↗
45

Octo: An Open-Source Generalist Robot Policy

Octo Model Team, Dibya Ghosh, Homer Walke +15 authors

Octo, a large transformer-based policy trained on extensive robotic datasets, demonstrates versatility and efficient fine-tuning for diverse robotic platforms and tasks.

29transformer-based policyOpen X-Embodiment datasetHF ↗arXiv ↗
46

Imp: Highly Capable Large Multimodal Models for Mobile Devices

Zhenwei Shao, Zhou Yu, Jun Yu +5 authors

A systematic study of lightweight large multimodal models (LMMs) led to the development of the Imp family, which achieves superior performance compared to larger models and is capable of high-speed inference on mobile devices.

29large language modelslarge multimodal modelsHF ↗arXiv ↗
47

Self-Play Preference Optimization for Language Model Alignment

Yue Wu, Zhiqing Sun, Huizhuo Yuan +3 authors

A self-play method called SPPO for language model alignment achieves state-of-the-art performance by approximating Nash equilibrium policy in a constant-sum game setting, outperforming other approaches with limited data.

29reinforcement learning from human feedbackparametric modelsHF ↗arXiv ↗
50

The Road Less Scheduled

Aaron Defazio, Xingyu, Yang +4 authors

A new schedule-free method achieves state-of-the-art performance in optimization without requiring a stopping time or additional hyper-parameters, by unifying scheduling and iterate averaging.

27learning rate schedulesoptimization stopping stepHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号