TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Mar 2 – Mar 8, 2026
本周最热199

Heterogeneous Agent Collaborative Reinforcement Learning

Zhixia Zhang, Zixuan Huang, Xin Xia +7 authors

HACRL enables collaborative reinforcement learning where heterogeneous agents share verified rollouts during training to improve collectively while maintaining independent operation at inference time, with HACPO achieving superior performance through efficient sample utilization and cross-agent knowledge transfer.

heterogeneous agentscollaborative optimizationon-policy optimizationmulti-agent reinforcement learningHF ↗arXiv ↗

50 篇论文 · 按点赞排序

02

Helios: Real Real-Time Long Video Generation Model

Shenghai Yuan, Yuanyang Yin, Zongjian Li +3 authors

Helios is a 14 billion parameter autoregressive diffusion model for video generation that achieves real-time performance and high-quality long-video synthesis without conventional optimization techniques.

190autoregressive diffusion modelvideo generationHF ↗arXiv ↗
03

Utonia: Toward One Encoder for All Point Clouds

Yujia Zhang, Xiaoyang Wu, Yunhan Yang +6 authors

Utonia enables cross-domain point cloud representation learning through a unified self-supervised transformer encoder, enhancing perception and supporting embodied and multimodal reasoning tasks.

187point transformer encoderself-supervised learningHF ↗arXiv ↗
04

dLLM: Simple Diffusion Language Modeling

Zhanhui Zhou, Lingjie Chen, Hanghang Tong +1 authors

A unified open-source framework is presented that standardizes core components of diffusion language modeling for reproduction, customization, and accessible development of both large and small models.

154diffusion language modelstrainingHF ↗arXiv ↗
10

SkillNet: Create, Evaluate, and Connect AI Skills

Yuan Liang, Ruobin Zhong, Haoming Xu +46 authors

SkillNet introduces an open infrastructure for systematically accumulating and transferring AI skills through a unified ontology, significantly improving agent performance across multiple domains.

95AI agentsskill consolidationHF ↗arXiv ↗
11

SWE-rebench V2: Language-Agnostic SWE Task Collection at Scale

Ibragim Badertdinov, Maksim Nekrashevich, Anton Shevtsov +1 authors

A large-scale dataset of software engineering tasks spanning multiple programming languages and repositories was created using an automated pipeline that generates executable environments and filters unreliable instances through LLM validation.

92reinforcement learningsoftware engineering agentsHF ↗arXiv ↗
13

UniG2U-Bench: Do Unified Models Advance Multimodal Understanding?

Zimo Wen, Boxiu Li, Wanbo Zhang +11 authors

Unified multimodal models show mixed performance in generation-to-understanding tasks, with specific subtasks benefiting from enhanced spatial and reasoning capabilities while overall performance lags behind specialized vision-language models.

88Unified multimodal modelsVision-Language ModelsHF ↗arXiv ↗
14

Qwen3-Coder-Next Technical Report

Ruisheng Cao, Mouxiang Chen, Jiawei Chen +17 authors

Qwen3-Coder-Next is an 80-billion-parameter language model that activates only 3 billion parameters during inference, achieving strong coding capabilities through agentic training with verifiable task synthesis and reinforcement learning.

66language modelparameter-efficient fine-tuningHF ↗arXiv ↗
18

CHIMERA: Compact Synthetic Data for Generalizable LLM Reasoning

Xinyu Zhu, Yihao Feng, Yanchao Sun +5 authors

A synthetic reasoning dataset called CHIMERA is introduced to overcome data-centric challenges in training large language models for cross-domain reasoning, achieving performance comparable to much larger models.

57Chain-of-Thoughtsupervised fine-tuningHF ↗arXiv ↗
19

DARE: Aligning LLM Agents with the R Statistical Ecosystem via Distribution-Aware Retrieval

Maojun Sun, Yue Wu, Yifei Xie +5 authors

A lightweight retrieval model called DARE incorporates data distribution information into function representations to improve R package retrieval, achieving superior performance over existing embedding models while enabling more reliable statistical analysis through an R-oriented LLM agent.

54retrieval-augmented approachesfunction-level semanticsHF ↗arXiv ↗
20

OpenAutoNLU: Open Source AutoML Library for NLU

Grigory Arshinov, Aleksandr Boriskin, Sergey Senichev +4 authors

OpenAutoNLU is an open-source automated machine learning library for NLU tasks that employs data-aware training selection and includes integrated diagnostics and LLM features through a minimal low-code interface.

50automated machine learningnatural language understandingHF ↗arXiv ↗
23

Mode Seeking meets Mean Seeking for Fast Long Video Generation

Shengqu Cai, Weili Nie, Chao Liu +8 authors

A training paradigm combining mode seeking and mean seeking in a Decoupled Diffusion Transformer enables efficient generation of high-quality long videos by leveraging both global flow matching and local distribution matching techniques.

41Decoupled Diffusion Transformerflow matchingHF ↗arXiv ↗
24

VGGT-Det: Mining VGGT Internal Priors for Sensor-Geometry-Free Multi-View Indoor 3D Object Detection

Yang Cao, Feize Wu, Dave Zhenyu Chen +3 authors

VGGT-Det enables sensor-geometry-free multi-view indoor 3D object detection by integrating a Visual Geometry Grounded Transformer encoder with attention-guided query generation and query-driven feature aggregation to effectively utilize semantic and geometric priors.

40Visual Geometry Grounded Transformermulti-view indoor 3D object detectionHF ↗arXiv ↗
27

RoboPocket: Improve Robot Policies Instantly with Your Phone

Junjie Fang, Wendi Chen, Han Xue +7 authors

RoboPocket enables efficient robot-free policy iteration using smartphone-based augmented reality visualization and asynchronous online fine-tuning to improve data and sample efficiency in imitation learning.

35imitation learningdata collectionHF ↗arXiv ↗
29

MemSifter: Offloading LLM Memory Retrieval via Outcome-Driven Proxy Reasoning

Jiejun Tan, Zhicheng Dou, Liancheng Zhang +3 authors

MemSifter is a framework that uses a small proxy model to offload memory retrieval from large language models, employing reinforcement learning with task-performance rewards and training techniques like curriculum learning and model merging to improve long-term memory efficiency and accuracy.

32Large Language Modelslong-term memoryHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号