TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

686 篇论文 · 按点赞排序

212

δ-mem: Efficient Online Memory for Large Language Models

Jingdi Lei, Di Zhang, Junxian Li +7 authors

A lightweight memory mechanism called δ-mem enhances large language models by augmenting a frozen attention backbone with a compact associative memory state that provides low-rank corrections to attention computations.

131large language modelsmemory mechanismHF ↗arXiv ↗
214

Are We Ready For An Agent-Native Memory System?

Wei Zhou, Xuanhe Zhou, Shaokun Han +5 authors

Large language model agents' memory systems have evolved into complex data management frameworks requiring systematic evaluation across multiple modules and workloads to understand their performance characteristics and trade-offs.

130large language model agentsmemory representation and storageHF ↗arXiv ↗
218

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

Team HY-World, Chenjie Cao, Xuhui Zuo +42 authors

HY-World 2.0 is a multi-modal world model framework that generates high-fidelity 3D Gaussian Splatting scenes from diverse inputs using specialized modules for panorama generation, trajectory planning, world expansion, and composition, along with an enhanced rendering platform for interactive 3D exploration.

127multi-modal world model3D Gaussian SplattingHF ↗arXiv ↗
220

RLDX-1 Technical Report

Dongyoung Kim, Huiwon Jang, Myungkyu Koo +65 authors

RLDX-1 is a general-purpose robotic policy for dexterous manipulation that integrates heterogeneous modalities through a Multi-Stream Action Transformer architecture, demonstrating superior performance in complex real-world tasks compared to existing vision-language-action models.

126Vision-Language-Action modelsMulti-Stream Action TransformerHF ↗arXiv ↗
222

daVinci-Dev: Agent-native Mid-training for Software Engineering

Ji Zeng, Dayuan Fu, Tiantian Mi +14 authors

Agentic mid-training enables large language models to develop autonomous software engineering capabilities through specialized data synthesis techniques that bridge the gap between static training data and dynamic development environments.

126Large Language Modelagentic software engineeringHF ↗arXiv ↗
226

Omni Interaction Agent Technical Report

Orantqing, Shengpeng Ji, Junlong Tong +20 authors

Gander is an end-to-end framework that integrates continuous multi-modal streaming, real-time full-duplex interaction, and agentic reasoning through a Cerebellum-Brain architecture and a chunk-level token stream design.

125Cerebellum-Brain collaborative frameworkstreaming Thinker-Talker architectureHF ↗arXiv ↗
234

SWE-Explore: Benchmarking How Coding Agents Explore Repositories

Shaoqiu Zhang, Yuhang Wang, Jialiang Liang +8 authors

SWE-Explore introduces a benchmark for evaluating coding agents' repository exploration capabilities by requiring ranked lists of relevant code regions within line budgets, demonstrating that agentic exploration outperforms traditional retrieval methods.

123repository explorationcoding agentsHF ↗arXiv ↗
236

MMSkills: Towards Multimodal Skills for General Visual Agents

Kangning Zhang, Shuai Shao, Qingyao Li +8 authors

Multimodal procedural knowledge frameworks enable visual agents to leverage external reusable skills through structured representations combining text, state cards, and visual keyframes, improving decision-making in complex environments.

122multimodal procedural knowledgevisual agentsHF ↗arXiv ↗
238

Audio Interaction Model

Zhifei Xie, Zihang Liu, Ze An +8 authors

A unified streaming audio model is developed that combines offline task execution with real-time audio instruction following through an end-to-end framework supporting multiple audio interaction capabilities.

121Large Audio Language Modelsstreaming audio modelsHF ↗arXiv ↗
8 / 23

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号