TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

本年最热779

BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

Björn Engdahl, Adrian Kosowski, Jan Chorowski +6 authors

A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.

in-context learningrecurrent latent reasoninghigh-dimensional latent spaceARC-AGI-1HF ↗arXiv ↗

682 篇论文 · 按点赞排序

04

A Very Big Video Reasoning Suite

Maijunxian Wang, Ruisi Wang, Juyi Lin +53 authors

A large-scale video reasoning dataset and benchmark are introduced to study video intelligence capabilities beyond visual quality, enabling systematic analysis of spatiotemporal reasoning and generalization across diverse tasks.

526video reasoningspatiotemporal consistencyHF ↗arXiv ↗
05

Kimi K3: Open Frontier Intelligence

Kimi Team, Tongtong Bai, Yifan Bai +399 authors

Kimi K3 is a large-scale mixture-of-experts model with native vision and long-context capabilities that improves scaling efficiency and achieves strong performance across coding, reasoning, and agentic tasks.

509Mixture-of-ExpertsKimi Delta AttentionHF ↗arXiv ↗
07

Orca: The World is in Your Mind

Yihao Wang, Yuheng Ji, Mingyu Cao +54 authors

Orca establishes a unified world latent space through next-state-prediction modeling using multimodal data and demonstrates superior performance in downstream tasks compared to specialized baselines.

506world foundation modelworld latent spaceHF ↗arXiv ↗
08

ABot-Earth 0.5: Generative 3D Earth Model

Ming Qian, Tianjian Ouyang, Mingchao Sun +25 authors

ABot-Earth 0.5 generates realistic 3D environments from satellite imagery using 3D Gaussian Splatting representation, enabling fast synthesis and real-time visualization for Embodied AI applications.

4873D Gaussian Splattinggenerative modelHF ↗arXiv ↗
09

StudentSim: Training LLM-based Student Simulators

Ke Yang, Chenglong Wang, Michel Galley +4 authors

StudentSim trains personalized student simulators from sparse data to mirror learner responses and adapt to tutor guidance, outperforming existing models across chess, writing, and math.

486student simulatorspooled trainingHF ↗arXiv ↗
10

Looped World Models

Hongyuan Adam Lu, Z. L. Victor Wei, Qun Zhang +28 authors

Looped World Models introduce iterative latent state refinement through shared transformer blocks, achieving 100x parameter efficiency while adapting computational depth to prediction complexity.

483world modelslooped architecturesHF ↗arXiv ↗
13

AI Can Learn Scientific Taste

Jingqi Tong, Mingzhe Li, Hangcheng Li +20 authors

Great scientists have strong judgement and foresight, closely tied to what we call scientific taste. Here, we use the term to refer to the capacity to judge and propose research ideas with high potential impact. However, most relative research focuses on improving an AI scientist's executive capability, while enhancing an AI's scientific taste remains underexplored. In this work, we propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference modeling and alignment problem. For preference modeling, we train Scientific Judge on 700K field- and time-matched pairs of high- vs. low-citation papers to judge ideas. For preference alignment, using Scientific Judge as a reward model, we train a policy model, Scientific Thinker, to propose research ideas with high potential impact. Experiments show Scientific Judge outperforms SOTA LLMs (e.g., GPT-5.2, Gemini 3 Pro) and generalizes to future-year test, unseen fields, and peer-review preference. Furthermore, Scientific Thinker proposes research ideas with higher potential impact than baselines. Our findings show that AI can learn scientific taste, marking a key step toward reaching human-level AI scientists.

431HF ↗arXiv ↗
14

Scaling Automatic Research Agents via World Models

Xiyuan Yang, Sheikh Sarwar, Jingru Cheng +8 authors

World Model RL replaces costly environment execution with a learned world model and applies debiasing and denoising to accelerate post-training of autonomous research agents.

422World Model RLAutoResearchHF ↗arXiv ↗
16

Agents' Last Exam

Yiyou Sun, Xinyang Han, Weichen Zhang +305 authors

Agents' Last Exam (ALE) is a benchmark for evaluating AI agents on long-term, economically valuable real-world tasks across 13 industry clusters with 1K+ tasks, revealing significant gaps between benchmark performance and practical deployment.

385AI agentsbenchmarkHF ↗arXiv ↗
19

Demystifing Video Reasoning

Ruisi Wang, Zhongang Cai, Fanyi Pu +11 authors

Diffusion-based video models demonstrate reasoning capabilities through denoising steps rather than frame sequences, exhibiting behaviors like working memory, self-correction, and perception-before-action within specialized transformer layers.

373diffusion modelsvideo generationHF ↗arXiv ↗
21

MolmoAct2: Action Reasoning Models for Real-world Deployment

Haoquan Fang, Jiafei Duan, Donovan Clay +26 authors

MolmoAct2 presents an open-action reasoning model for robotics that improves upon previous systems through specialized vision-language-model backbones, new datasets, open-weight action tokenizers, architectural redesign for continuous-action prediction, and adaptive reasoning for reduced latency.

355Vision-Language-Action modelsVLM backboneHF ↗arXiv ↗
29

HRM-Text: Efficient Pretraining Beyond Scaling

Guan Wang, Changling Liu, Chenyu Wang +6 authors

A Hierarchical Recurrent Model architecture with specialized training on instruction-response pairs achieves competitive language modeling performance with significantly reduced computational requirements compared to traditional Transformer-based approaches.

321Hierarchical Recurrent ModelTransformersHF ↗arXiv ↗
1 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号