TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 22 – Sep 28, 2025

50 篇论文 · 按点赞排序

32

Thinking Augmented Pre-training

Liang Wang, Nan Yang, Shaohan Huang +2 authors

Thinking augmented pre-training improves data efficiency and performance of large language models by augmenting text with automatically generated thinking trajectories.

24large language modelthinking trajectoriesHF ↗arXiv ↗
39

BaseReward: A Strong Baseline for Multimodal Reward Model

Yi-Fan Zhang, Haihua Yang, Huanyu Zhang +12 authors

The paper provides a comprehensive guide and introduces BaseReward, a state-of-the-art multimodal reward model, which outperforms existing models across various benchmarks and real-world tasks.

21Multimodal Large Language ModelsReward ModelsHF ↗arXiv ↗
49

Soft Tokens, Hard Truths

Natasha Butt, Ariel Kwiatkowski, Ismail Labiad +2 authors

A scalable reinforcement learning method for learning continuous chain-of-thought tokens in large language models improves performance and diversity over discrete tokens while preserving out-of-domain predictions.

16Chain-of-Thought (CoT)continuous tokensHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号