TensorX
返回文献探索

Paper · arXiv 2305.02968

Masked Trajectory Models for Prediction, Representation, and Control

Philipp Wu, Arjun Majumdar, Kevin Stone, Yixin Lin, Igor Mordatch, Pieter Abbeel, Aravind Rajeswaran

1 upvotesMay 4, 2023arXiv 预印本
AI 摘要

Masked Trajectory Models (MTM) are versatile sequential decision-making models that can perform tasks such as forward dynamics modeling, inverse dynamics modeling, and offline RL using generic masks and representations, matching or outperforming specialized methods.

Masked Trajectory ModelsMTMsequential decision makingtrajectorystate-action sequenceforward dynamics modelinverse dynamics modeloffline RL agentstate representationstraditional RL algorithms

Abstract

We introduce Masked Trajectory Models (MTM) as a generic abstraction for sequential decision making. MTM takes a trajectory, such as a state-action sequence, and aims to reconstruct the trajectory conditioned on random subsets of the same trajectory. By training with a highly randomized masking pattern, MTM learns versatile networks that can take on different roles or capabilities, by simply choosing appropriate masks at inference time. For example, the same MTM network can be used as a forward dynamics model, inverse dynamics model, or even an offline RL agent. Through extensive experiments in several continuous control tasks, we show that the same MTM network -- i.e. same weights -- can match or outperform specialized networks trained for the aforementioned capabilities. Additionally, we find that state representations learned by MTM can significantly accelerate the learning speed of traditional RL algorithms. Finally, in offline RL benchmarks, we find that MTM is competitive with specialized offline RL algorithms, despite MTM being a generic self-supervised learning method without any explicit RL components. Code is available at https://github.com/facebookresearch/mtm

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号