TensorX
返回文献探索

Paper · arXiv 2509.18644

Do You Need Proprioceptive States in Visuomotor Policies?

Juntu Zhao, Wenbo Lu, Di Zhang, Yufeng Liu, Yushen Liang, Tianluo Zhang, Yifeng Cao, Junyuan Xie, Yingdong Hu, Shengjie Wang, Junliang Guo, Dequan Wang, Yang Gao

50 upvotesSeptember 23, 2025arXiv 预印本
AI 摘要

A state-free policy using only visual observations achieves better spatial generalization and data efficiency in robot manipulation tasks compared to state-based policies.

imitation-learning-based visuomotor policiesproprioceptive state inputoverfittingspatial generalizationrelative end-effector action spacedual wide-angle wrist cameraspick-and-placeshirt-foldingwhole-body manipulationcross-embodiment adaptation

Abstract

Imitation-learning-based visuomotor policies have been widely used in robot manipulation, where both visual observations and proprioceptive states are typically adopted together for precise control. However, in this study, we find that this common practice makes the policy overly reliant on the proprioceptive state input, which causes overfitting to the training trajectories and results in poor spatial generalization. On the contrary, we propose the State-free Policy, removing the proprioceptive state input and predicting actions only conditioned on visual observations. The State-free Policy is built in the relative end-effector action space, and should ensure the full task-relevant visual observations, here provided by dual wide-angle wrist cameras. Empirical results demonstrate that the State-free policy achieves significantly stronger spatial generalization than the state-based policy: in real-world tasks such as pick-and-place, challenging shirt-folding, and complex whole-body manipulation, spanning multiple robot embodiments, the average success rate improves from 0\% to 85\% in height generalization and from 6\% to 64\% in horizontal generalization. Furthermore, they also show advantages in data efficiency and cross-embodiment adaptation, enhancing their practicality for real-world deployment.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Do You Need Proprioceptive States in Visuomotor Policies? | TensorX