TensorX
返回文献探索

Paper · arXiv 2310.16828

TD-MPC2: Scalable, Robust World Models for Continuous Control

Nicklas Hansen, Hao Su, Xiaolong Wang

8 upvotesOctober 25, 2023arXiv 预印本
AI 摘要

TD-MPC2, an improved model-based reinforcement learning algorithm, achieves strong performance across various RL tasks with consistent results using a single set of hyperparameters and benefits from larger models and datasets.

model-based reinforcement learningRLTD-MPCTD-MPC2local trajectory optimizationlatent spaceimplicit world modeldecoder-freehyperparameters

Abstract

TD-MPC is a model-based reinforcement learning (RL) algorithm that performs local trajectory optimization in the latent space of a learned implicit (decoder-free) world model. In this work, we present TD-MPC2: a series of improvements upon the TD-MPC algorithm. We demonstrate that TD-MPC2 improves significantly over baselines across 104 online RL tasks spanning 4 diverse task domains, achieving consistently strong results with a single set of hyperparameters. We further show that agent capabilities increase with model and data size, and successfully train a single 317M parameter agent to perform 80 tasks across multiple task domains, embodiments, and action spaces. We conclude with an account of lessons, opportunities, and risks associated with large TD-MPC2 agents. Explore videos, models, data, code, and more at https://nicklashansen.github.io/td-mpc2

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
TD-MPC2: Scalable, Robust World Models for Continuous Control | TensorX