TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Sep 29 – Oct 5, 2025
本周最热555

The Dragon Hatchling: The Missing Link between the Transformer and Models of the Brain

Adrian Kosowski, Przemysław Uznański, Jan Chorowski +2 authors

BDH, a biologically inspired Large Language Model, combines scale-free network architecture with Hebbian learning to achieve Transformer-like performance while maintaining interpretability.

Large Language Modelscale-free networkHebbian learningsynaptic plasticityHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

LongLive: Real-time Interactive Long Video Generation

Shuai Yang, Wei Huang, Ruihang Chu +9 authors

LongLive is a frame-level autoregressive framework for real-time and interactive long video generation, addressing efficiency and quality challenges through causal attention, KV-recache, streaming long tuning, and short window attention.

191frame-level autoregressivediffusion modelsHF ↗arXiv ↗
09

Quantile Advantage Estimation for Entropy-Safe Reasoning

Junkang Wu, Kexin Huang, Jiancan Wu +3 authors

Quantile Advantage Estimation stabilizes reinforcement learning with verifiable rewards by addressing entropy issues and improving performance on large language models.

119Reinforcement Learning with Verifiable Rewardsentropy collapseHF ↗arXiv ↗
12

GEM: A Gym for Agentic LLMs

Zichen Liu, Anya Sims, Keyu Duan +16 authors

GEM, an open-source environment simulator, facilitates experience-based learning for large language models by providing a standardized framework and diverse environments for training and benchmarking reinforcement learning algorithms.

92large language modelsexperience-based learningHF ↗arXiv ↗
13

ExGRPO: Learning to Reason from Experience

Runzhe Zhan, Yafu Li, Zhi Wang +5 authors

ExGRPO, a framework that prioritizes valuable reasoning experiences, improves and stabilizes reinforcement learning from verifiable rewards for large language models.

83reinforcement learning from verifiable rewardsRLVRHF ↗arXiv ↗
17

Variational Reasoning for Language Models

Xiangxin Zhou, Zichen Liu, Haonan Wang +5 authors

A variational reasoning framework treats thinking traces as latent variables, optimizing them through variational inference to improve language model reasoning.

70variational reasoning frameworklatent variablesHF ↗arXiv ↗
22

Multiplayer Nash Preference Optimization

Fang Wu, Xu Huang, Weihao Xuan +8 authors

Multiplayer Nash Preference Optimization (MNPO) extends Nash learning from human feedback to handle complex, non-transitive human preferences by formulating alignment as an n-player game.

62Reinforcement learning from human feedbackRLHFHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号