TensorX
返回文献探索

Paper · arXiv 2306.10007

Robot Learning with Sensorimotor Pre-training

Ilija Radosavovic, Baifeng Shi, Letian Fu, Ken Goldberg, Trevor Darrell, Jitendra Malik

14 upvotesJune 16, 2023arXiv 预印本
AI 摘要

A self-supervised sensorimotor pre-training method using a Transformer model on sequences of sensorimotor tokens improves robotic performance in tasks like block stacking.

Transformersensorimotor tokenscamera imagesproprioceptive statespast actionslatent visual representations

Abstract

We present a self-supervised sensorimotor pre-training approach for robotics. Our model, called RPT, is a Transformer that operates on sequences of sensorimotor tokens. Given a sequence of camera images, proprioceptive robot states, and past actions, we encode the interleaved sequence into tokens, mask out a random subset, and train a model to predict the masked-out content. We hypothesize that if the robot can predict the missing content it has acquired a good model of the physical world that can enable it to act. RPT is designed to operate on latent visual representations which makes prediction tractable, enables scaling to 10x larger models, and 10 Hz inference on a real robot. To evaluate our approach, we collect a dataset of 20,000 real-world trajectories over 9 months using a combination of motion planning and model-based grasping algorithms. We find that pre-training on this data consistently outperforms training from scratch, leads to 2x improvements in the block stacking task, and has favorable scaling properties.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Robot Learning with Sensorimotor Pre-training | TensorX