TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jul 27 – Aug 2, 2026

50 篇论文 · 按点赞排序

34

Data Pyramid for Embodied Manipulation

Yifan Ye, Yankai Fu, Yaoxu Lv +26 authors

The paper organizes embodied robot learning data into a scalable pyramid and analyzes how mixing sources affects foundation model capabilities and open challenges.

36embodied foundation modelsvision-language-action modelsHF ↗arXiv ↗
36

Pass the Baton: Trajectory-Relayed On-Policy Distillation

Haolei Xu, Xiaowen Xu, Haiwen Hong +5 authors

Relay-OPD detects reasoning errors via teacher-student trajectory divergence and briefly delegates generation to the teacher to recover valid supervision, improving small-model math reasoning while cutting training cost.

33on-policy distillationprefix failureHF ↗arXiv ↗
39

RefCaptioner: Multi-Reference Image-Grounded Video Captioning

Tengfei Liu, Yang Shi, Yuran Wang +16 authors

RefCaptioner is a two-stage framework for multi-reference image-grounded video captioning that improves phrase-level grounding and factual consistency, supported by a new benchmark and training corpus.

30multi-reference image-grounded video captioningRefCaptionerHF ↗arXiv ↗
42

Scaling Native Multimodal Pre-Training From Scratch

Haoyuan Wu, Aoqi Wu, Hai Wang +3 authors

Native multimodal pre-training of vision-language transformers follows predictable compute scaling laws with distinct language and multimodal allocation behaviors, enabling efficient resource allocation and cross-modal transfer.

29native multimodal pre-trainingvision-language modelHF ↗arXiv ↗
48

See2Think: Do Multimodal Models Really Use Intermediate Visual States?

Siyu Yan, Zhuoran Yan, Haiying Xu +10 authors

See2Think evaluates whether multimodal language models genuinely use intermediate visual reasoning states through a benchmark of visually dependent tasks and a process-tracing framework, revealing that faithful rendering is a key bottleneck and that models depend behaviorally on visual feedback.

25multimodal large language modelsvisual reasoningHF ↗arXiv ↗
50

Shieldstral

Antonia Calvi, Avinash Sooriyarachchi, Giada Pistilli +8 authors

Shieldstral is a compact multimodal safety classifier that unifies diverse moderation tasks as binary questions to achieve high performance with far fewer parameters.

24policy-adaptive multimodal safety classifierbinary question-answeringHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号