TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 8 – Sep 14, 2025
本周最热664

Sharing is Caring: Efficient LM Post-Training with Collective RL Experience Sharing

Jeffrey Amico, Gabriel Passamani Andrade, John Donaghy +12 authors

Swarm sAmpling Policy Optimization (SAPO) is a decentralized and asynchronous RL algorithm that enhances post-training language models without supervised fine-tuning, achieving significant reward gains and scalability across diverse hardware.

reinforcement learningpost-traininglanguage modelsDeepSeek-R1-ZeroHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

Why Language Models Hallucinate

Adam Tauman Kalai, Ofir Nachum, Santosh S. Vempala +1 authors

Language models produce incorrect statements due to training and evaluation procedures that reward guessing over acknowledging uncertainty, leading to a need for socio-technical changes in benchmark scoring.

200hallucinationsbinary classificationHF ↗arXiv ↗
07

Parallel-R1: Towards Parallel Thinking via Reinforcement Learning

Tong Zheng, Hongming Zhang, Wenhao Yu +8 authors

Parallel-R1, a reinforcement learning framework, enhances large language models' reasoning capabilities by enabling parallel thinking through a progressive curriculum, leading to significant performance improvements on math benchmarks.

106parallel thinkingreinforcement learningHF ↗arXiv ↗
11

RewardDance: Reward Scaling in Visual Generation

Jie Wu, Yu Gao, Zilyu Ye +9 authors

RewardDance is a scalable reward modeling framework that aligns with VLM architectures, enabling effective scaling of RMs and resolving reward hacking issues in generation models.

73CLIP-based RMsBradley-Terry lossesHF ↗arXiv ↗
15

3D and 4D World Modeling: A Survey

Lingdong Kong, Wesley Yang, Jianbiao Mei +20 authors

This survey provides a comprehensive review of 3D and 4D world modeling and generation, establishing definitions, taxonomy, datasets, and evaluation metrics, and discussing applications and challenges.

59world models3D world modelingHF ↗arXiv ↗
18

Set Block Decoding is a Language Model Inference Accelerator

Itai Gat, Heli Ben-Hamu, Marton Havasi +6 authors

Set Block Decoding accelerates language model generation by integrating next token prediction and masked token prediction, enabling parallel sampling of future tokens and reducing computational cost without sacrificing accuracy.

54autoregressive next token predictionmasked token predictionHF ↗arXiv ↗
24

Reconstruction Alignment Improves Unified Multimodal Models

Ji Xie, Trevor Darrell, Luke Zettlemoyer +1 authors

Reconstruction Alignment (RecA) is a post-training method that enhances multimodal models by using visual embeddings as dense prompts, improving image generation and editing fidelity.

40Unified multimodal modelsReconstruction AlignmentHF ↗arXiv ↗
25

Does DINOv3 Set a New Medical Vision Standard?

Che Liu, Yinda Chen, Haoyuan Shi +16 authors

DINOv3, a self-supervised vision transformer, demonstrates strong performance across various medical vision tasks without domain-specific pre-training, though it shows limitations in deeply specialized domains and does not consistently follow scaling laws in the medical domain.

40self-supervised vision transformerViTHF ↗arXiv ↗
30

Language Self-Play For Data-Free Training

Jakub Grudzien Kuba, Mengting Gu, Qi Ma +2 authors

Language Self-Play (LSP) enhances large language models' performance on instruction-following tasks through self-play, surpassing data-driven methods.

32large language modelsreinforcement learningHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号