TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Apr 13 – Apr 19, 2026
本周最热248

WildDet3D: Scaling Promptable 3D Detection in the Wild

Weikai Huang, Jieyu Zhang, Sijun Li +14 authors

A unified 3D object detection framework with a large-scale dataset enables open-world detection with multiple prompt types and geometric cue integration.

monocular 3D object detectiongeometry-aware architecturetext promptspoint promptsHF ↗arXiv ↗

50 篇论文 · 按点赞排序

03

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping

Yang Liu, Enxi Wang, Yufei Gao +6 authors

MEDS is a memory-enhanced dynamic reward shaping framework that improves sampling diversity in reinforcement learning for large language models by identifying and penalizing recurrent error patterns through clustering of historical behavioral signals.

144reinforcement learninglarge language modelsHF ↗arXiv ↗
05

HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds

Team HY-World, Chenjie Cao, Xuhui Zuo +42 authors

HY-World 2.0 is a multi-modal world model framework that generates high-fidelity 3D Gaussian Splatting scenes from diverse inputs using specialized modules for panorama generation, trajectory planning, world expansion, and composition, along with an enhanced rendering platform for interactive 3D exploration.

127multi-modal world model3D Gaussian SplattingHF ↗arXiv ↗
08

Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe

Yaxuan Li, Yuxin Zuo, Bingxiang He +8 authors

On-policy distillation dynamics in large language models depend on compatible thinking patterns between teacher and student models, with successful distillation characterized by alignment on high-probability tokens and requiring teachers to provide novel capabilities beyond student training data.

116on-policy distillationlarge language modelsHF ↗arXiv ↗
11

FORGE:Fine-grained Multimodal Evaluation for Manufacturing Scenarios

Xiangru Jian, Hao Xu, Wei Pang +13 authors

FORGE introduces a high-quality multimodal manufacturing dataset with fine-grained domain semantics to evaluate MLLMs on real-world tasks, revealing that domain-specific knowledge rather than visual grounding limits performance, and demonstrating that supervised fine-tuning on structured annotations significantly improves accuracy.

99Multimodal Large Language Modelsvisual groundingHF ↗arXiv ↗
13

EXAONE 4.5 Technical Report

Eunbi Choi, Kibong Choi, Sehyun Chun +55 authors

EXAONE 4.5 is an open-weight vision language model that integrates a visual encoder into EXAONE 4.0, achieving enhanced document understanding and general language capabilities through targeted data curation and extended context length.

75vision language modelvisual encoderHF ↗arXiv ↗
22

Lyra 2.0: Explorable Generative 3D Worlds

Tianchang Shen, Sherwin Bahmani, Kai He +12 authors

Lyra 2.0 enables large-scale 3D scene creation through persistent video generation that addresses spatial forgetting and temporal drifting issues in long-horizon video models.

41video generation3D scene creationHF ↗arXiv ↗
24

Geometric Context Transformer for Streaming 3D Reconstruction

Lin-Zhuo Chen, Jian Gao, Yihang Chen +8 authors

LingBot-Map is a feed-forward 3D foundation model that reconstructs scenes from video streams using a geometric context transformer architecture with specialized attention mechanisms for coordinate grounding, dense geometric cues, and long-range drift correction, achieving stable real-time performance at 20 FPS.

39Simultaneous Localization and Mappingfeed-forward 3D foundation modelHF ↗arXiv ↗
26

CodeTracer: Towards Traceable Agent States

Han Li, Yifan Yao, Letian Zhu +13 authors

CodeTracer is a tracing architecture that analyzes code agent execution by reconstructing state transitions and localizing failures in complex multi-stage workflows.

38code agentsagent tracingHF ↗arXiv ↗
27

CocoaBench: Evaluating Unified Digital Agents in the Wild

CocoaBench Team, Shibo Hao, Zhining Zhang +29 authors

A new benchmark called CocoaBench evaluates unified digital agents on complex, multi-capability tasks requiring vision, search, and coding integration, revealing significant room for improvement in current agent systems.

37LLM agentssoftware engineeringHF ↗arXiv ↗
30

Toward Autonomous Long-Horizon Engineering for ML Research

Guoxin Chen, Jie Chen, Lei Chen +7 authors

AiScientist enables autonomous long-horizon ML research engineering by combining hierarchical orchestration with durable state management, achieving superior performance on benchmark tasks through structured coordination and persistent project artifacts.

34autonomous AI researchlong-horizon ML research engineeringHF ↗arXiv ↗
1 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号