TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 6 – Apr 12, 2026
本周最热641

GrandCode: Achieving Grandmaster Level in Competitive Programming via Agentic Reinforcement Learning

DeepReinforce Team, Xiaoya Li, Xiaofei Sun +4 authors

GrandCode is a multi-agent reinforcement learning system that outperforms human competitors in competitive programming challenges by orchestrating specialized agent modules and employing novel reward policy optimization techniques.

multi-agent RLreinforcement learningagentic GRPOpost-trainingHF ↗arXiv ↗

50 篇论文 · 按点赞排序

05

ClawBench: Can AI Agents Complete Everyday Online Tasks?

Yuxuan Zhang, Yubo Wang, Yipeng Zhu +18 authors

ClawBench presents a comprehensive evaluation framework with 153 real-world tasks across 144 platforms to test AI agents' ability to automate everyday online activities requiring complex multi-step workflows and document processing.

264AI agentsevaluation frameworkHF ↗arXiv ↗
06

InCoder-32B-Thinking: Industrial Code World Model for Thinking

Jian Yang, Wei Zhang, Jiajun Wu +22 authors

Industrial software development lacks expert reasoning traces for hardware constraints, so a model was trained on error-driven reasoning chains and domain-specific execution traces to generate high-quality code reasoning and performance.

239Error-driven Chain-of-Thoughtindustrial code world modelHF ↗arXiv ↗
10

Self-Distilled RLVR

Chenxu Yang, Chuanyu Qin, Qingyi Si +7 authors

RLSD combines reinforcement learning with verifiable rewards and self-distillation to achieve stable training with fine-grained updates and reliable policy direction from environmental feedback.

181on-policy distillationon-policy self-distillationHF ↗arXiv ↗
17

LPM 1.0: Video-based Character Performance Model

Ailing Zeng, Casper Yang, Chauncey Ge +22 authors

A large-scale multimodal model for real-time conversational character performance generation that maintains identity consistency while enabling interactive, infinite-length video synthesis.

82Diffusion Transformermultimodal conditioningHF ↗arXiv ↗
19

A Simple Baseline for Streaming Video Understanding

Yujiao Shen, Shulin Tian, Jingkang Yang +1 authors

A simple sliding-window approach using recent video frames outperforms complex memory-based streaming video understanding methods, revealing trade-offs between real-time perception and long-term memory capabilities.

73streaming video understandingmemory mechanismsHF ↗arXiv ↗
20

Learning to Retrieve from Agent Trajectories

Yuqi Zhou, Sunhao Dai, Changle Qu +3 authors

Retrieval models for agentic search should be trained directly from agent interaction data using a new paradigm that mines supervision from multi-step agent trajectories and incorporates relevance intensity through weighted optimization.

72learning-to-ranklarge language modelsHF ↗arXiv ↗
21

RAGEN-2: Reasoning Collapse in Agentic RL

Zihan Wang, Chi Gui, Xing Jin +13 authors

Research identifies template collapse in multi-turn LLM agents as a hidden failure mode undetectable by entropy, proposing mutual information proxies and SNR-aware filtering to improve reasoning quality and task performance.

70entropymutual informationHF ↗arXiv ↗
22

Memory Intelligence Agent

Jingyang Qiao, Weicheng Meng, Yu Cheng +6 authors

Memory Intelligence Agent framework integrates non-parametric and parametric memory systems with reinforcement learning to enable efficient reasoning and autonomous evolution in open-world environments.

58Memory Intelligence AgentManager-Planner-Executor architectureHF ↗arXiv ↗
23

DMax: Aggressive Parallel Decoding for dLLMs

Zigeng Chen, Gongfan Fang, Xinyin Ma +2 authors

DMax introduces a novel approach for efficient diffusion language models that reduces error accumulation during parallel decoding through self-refinement and unified training strategies.

54diffusion language modelsparallel decodingHF ↗arXiv ↗
25

ACES: Who Tests the Tests? Leave-One-Out AUC Consistency for Code Generation

Hui Sun, Yun-Ji Zhang, Zheng Xie +4 authors

Researchers address the challenge of selecting correct code candidates from LLM-generated outputs by developing ACES, a method that ranks tests based on their ability to distinguish correct from incorrect code through leave-one-out evaluation and AUC consistency scoring.

53LLM-generated codetest correctnessHF ↗arXiv ↗
30

Vanast: Virtual Try-On with Human Image Animation via Synthetic Triplet Supervision

Hyunsoo Cha, Wonjung Woo, Byungjun Kim +1 authors

Vanast is a unified framework that generates garment-transferred human animation videos by combining image-based virtual try-on and pose-driven animation in a single process, addressing issues like identity drift and garment distortion through triplet supervision and dual module architecture.

48virtual try-onpose-driven animationHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号