TensorX

Explore · 每周精选

发现最受关注的研究论文,追踪研究趋势,订阅感兴趣的期刊与关键词。

Jun 23 – Jun 29, 2025

50 篇论文 · 按点赞排序

33

Unified Vision-Language-Action Model

Yuqi Wang, Xinghang Li, Wenxuan Wang +5 authors

UniVLA is a multimodal VLA model that autoregressively processes vision, language, and action as token sequences, incorporating world modeling for effective long-horizon policy learning and achieving state-of-the-art results across simulation and real-world benchmarks.

28vision-language-action modelsVLAsHF ↗arXiv ↗
42

Learning to Skip the Middle Layers of Transformers

Tim Lawson, Laurence Aitchison

A novel conditional computation architecture for Transformers dynamically skips middle layers based on input and a gating mechanism, but does not outperform dense baselines in reducing computational cost or improving validation performance.

18conditional computationTransformersHF ↗arXiv ↗
44

RoboTwin 2.0: A Scalable Data Generator and Benchmark with Strong Domain Randomization for Robust Bimanual Robotic Manipulation

Tianxing Chen, Zanxin Chen, Baijun Chen +23 authors

RoboTwin 2.0 is a scalable simulation framework for bimanual robotic manipulation that uses expert data synthesis and structured domain randomization to generate diverse and realistic synthetic data, improving sim-to-real transfer and generalization.

18multimodal large language modelssimulation-in-the-loopHF ↗arXiv ↗
45

Arch-Router: Aligning LLM Routing with Human Preferences

Co Tran, Salman Paracha, Adil Hafeez +1 authors

A preference-aligned routing framework using a compact 1.5B model effectively matches queries to user-defined domains and action types, outperforming proprietary models in subjective evaluation criteria.

17large language modelsLLM routingHF ↗arXiv ↗
46

SAM4D: Segment Anything in Camera and LiDAR Streams

Jianyun Xu, Song Wang, Ziqian Ni +4 authors

SAM4D is a multi-modal and temporal foundation model for segmentation in autonomous driving using Unified Multi-modal Positional Encoding and Motion-aware Cross-modal Memory Attention, with a multi-modal automated data engine generating pseudo-labels.

17multi-modaltemporal foundation modelHF ↗arXiv ↗
2 / 2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号