TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 3 – Mar 9, 2025

50 篇论文 · 按点赞排序

34

LADDER: Self-Improving LLMs Through Recursive Problem Decomposition

Toby Simonds, Akira Yoshiyama

LADDER, a framework for autonomous difficulty-driven problem-solving, and TTRL, a test-time reinforcement learning approach, enhance Large Language Models' performance on mathematical integration tasks without human supervision.

21LADDERLearning through Autonomous Difficulty-Driven Example RecursionHF ↗arXiv ↗
36

ABC: Achieving Better Control of Multimodal Embeddings using VLMs

Benjamin Schneider, Florian Kerschbaum, Wenhu Chen

ABC is a multimodal embedding model using a vision-language backbone that achieves top performance on MSCOCO, classification, and VQA tasks, offering high-quality representations with flexible natural language control.

20visual embedding modelszero-shot tasksHF ↗arXiv ↗
37

Mask-DPO: Generalizable Fine-grained Factuality Alignment of LLMs

Yuzhe Gu, Wenwei Zhang, Chengqi Lyu +2 authors

Mask-DPO, a fine-grained factuality alignment method based on Direct Preference Optimization, improves the factuality of LLM responses by learning from factually correct sentences and generalizes to unseen questions and topics.

20Large language models (LLMs)hallucinationsHF ↗arXiv ↗
43

FuseChat-3.0: Preference Optimization Meets Heterogeneous Model Fusion

Ziyi Yang, Fanqi Wan, Longguang Zhong +3 authors

FuseChat-3.0, a suite of large language models, integrates diverse source models into smaller targets using supervised fine-tuning and Direct Preference Optimization, achieving significant performance improvements across various tasks.

15large language modelsheterogeneous source LLMsHF ↗arXiv ↗
44

Iterative Value Function Optimization for Guided Decoding

Zhenhua Liu, Lijun Li, Ruizhe Chen +4 authors

A novel framework, Iterative Value Function Optimization, improves value-guided decoding in reinforcement learning by enhancing value function estimation accuracy, aligning language models and reducing computational costs.

15Reinforcement Learning from Human Feedback (RLHF)value-guided methodsHF ↗arXiv ↗
45

Efficient Test-Time Scaling via Self-Calibration

Chengsong Huang, Langlin Huang, Jixuan Leng +2 authors

Self-Calibration enhances the efficiency of test-time scaling in Large Language Models by using model confidence to dynamically adjust sampling strategies.

15Large Language Models (LLMs)Best-of-N samplingHF ↗arXiv ↗
49

Unified Video Action Model

Shuang Li, Yihuai Gao, Dorsa Sadigh +1 authors

The Unified Video Action (UVA) model integrates video and action predictions through a joint latent representation and decoupled decoding for high accuracy and efficient action inference in robotics.

14Unified Video Action modelUVAHF ↗arXiv ↗
50

Large-Scale Data Selection for Instruction Tuning

Hamish Ivison, Muru Zhang, Faeze Brahman +2 authors

A study on the scalability of data selection methods in instruction-tuning large language models reveals that a simpler, compute-efficient variant of representation-based data selection outperforms more complex methods.

14instruction-tuninglanguage modelsHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号