TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

September 2026

50 篇论文 · 按点赞排序

36

Language Models Can Control Their Own Attention

Namgyu Ho, Huzama Ahmad, Woosung Koh +3 authors

Declarative Attention lets language models declare relevant context regions during reasoning to skip most KV cache reads, reducing attended tokens with small accuracy trade-offs.

72Declarative AttentionKV cacheHF ↗arXiv ↗
40

Iris: Climbing to the Search Frontier

Ziyuan Liu, Hengqi Liu, Zichuan Wang +6 authors

Two large-scale search agents are trained via a multi-stage pipeline combining supervised fine-tuning and reinforcement learning against live search, achieving state-of-the-art open-source results on complex web benchmarks through rigorous trajectory filtering and inference-time context management.

64search agentsmulti-hop chainsHF ↗arXiv ↗
41

UI-Venus-2 Technical Report

Venus Team, Zhuohan Cai, Haoxing Chen +28 authors

UI-Venus-2 is a general-purpose multimodal GUI agent that uses unified reasoning-action loops, expanded environment coverage, and robust verification to enable reliable real-world digital automation.

64multimodal GUI agentsclosed-loop reasoning-action frameworkHF ↗arXiv ↗
43

Marigold V2: Revisiting Diffusion Transformers for Monocular Depth Estimation

Igor Pavlovic, Thiemo Wandel, Anton Obukhov +6 authors

Marigold V2 repurposes diffusion transformers for monocular depth estimation via single-step flow-matching inference, semantic alignment, and a Sinkhorn-based two-stage fine-tuning protocol, yielding sharper out-of-distribution depth maps and strong results on related dense regression tasks.

56monocular depth estimationdiffusion transformerHF ↗arXiv ↗
44

DriveZero: End-to-End Driving Beyond Human Demonstrations

Hao He, Chengcheng Hu, Zirun Su +17 authors

DriveZero is an end-to-end autonomous driving system that combines a vision foundation model for perception with a closed-loop reinforcement learning action model to learn driving behaviors beyond human demonstrations.

56DriveZeroDriveRLHF ↗arXiv ↗
45

Miles v0.1: Production-Level Post-Training

RadixArk, Tom Chen, Mao Cheng +10 authors

Miles is an open-source, production-ready system for large-scale reinforcement learning and post-training that supports diverse backends, weight synchronization, LoRA, distillation, and diffusion models.

52reinforcement-learningrollout enginesHF ↗arXiv ↗
46

GE-Act 2.0: Pretraining and Scaling a World-Action Model for Robotic Manipulation

AgiBot Research Team, Renhang Liu, Wenzhi Zhao +42 authors

GE-Act 2.0 is a world-action model trained from scratch with a control-oriented autoencoder, single-step visual planner, and inverse dynamics model, using knowledge-aligned selective optimization to enable scalable zero-shot robot manipulation across diverse skills and conditions.

52world-action modelscontrol-oriented autoencoderHF ↗arXiv ↗
47

Normalized Low-Rank Adaptation

Jiale Kang, Ziyin Yue, Zheng Zhan +2 authors

Normalized Low-Rank Adaptation stabilizes LoRA training by normalizing down-projection matrices, accelerating convergence and improving performance without extra parameters or inference cost.

52low-rank adaptationLoRAHF ↗arXiv ↗
49

H3-World: Turning Language Understanding into World Control

Danze Chen, Zeqing Wang, Ziyue Lin +2 authors

We present H3-World, an efficient framework that turns the 33B MiniMax-H3 video generator into an interactive world model. Our key finding is that, as large video generators become more capable, language is emerging as a natural interface for control. MiniMax-H3, for example, already supports zero-shot control of character behavior and camera motion through natural-language instructions. Building on this, H3-World turns this coarse language interface into precise, temporally grounded world control, without introducing dedicated action modules. Specifically, we represent each action as a structured combination of character and camera instructions, and align them with the corresponding temporal video latents. To make the control temporally precise, we further introduce temporal attention routing, which restricts each instruction to its intended time interval and reduces control leakage across actions. Importantly, H3-World directly reuses the semantic representations learned during large-scale video pretraining and requires only lightweight adaptation. With only 8,000 gameplay samples, 10,000 LoRA optimization steps, and 0.199% trainable parameters, H3-World achieves effective character and camera control while preserving strong generation quality. It also generalizes to unseen scenarios. These results show that the control capabilities emerging in large video generators can be efficiently transformed into interactive world control.

50HF ↗arXiv ↗
50

T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks

Junyao Yang, Yucheng Shi, Zhongzhi Li +4 authors

T1 is a 122B Mixture-of-Experts model trained with reinforcement learning to execute long-horizon terminal tasks in a cloud sandbox, achieving state-of-the-art results through stable actor-critic optimization and out-of-distribution training.

47Mixture-of-Expertsreinforcement learningHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号