TensorX
返回文献探索

Paper · arXiv 2309.08587

Compositional Foundation Models for Hierarchical Planning

Anurag Ajay, Seungwook Han, Yilun Du, Shaung Li, Abhi Gupta, Tommi Jaakkola, Josh Tenenbaum, Leslie Kaelbling, Akash Srivastava, Pulkit Agrawal

10 upvotesSeptember 15, 2023arXiv 预印本
AI 摘要

HiP, a compositional foundation model, integrates language, vision, and action models to plan and execute complex, long-horizon manipulation tasks through hierarchical reasoning and iterative refinement.

hierarchical reasoningsubgoal sequencessymbolic plansvideo diffusion modelinverse dynamics modelcompositional foundation modelsHierarchical Planning (HiP)long-horizon tasksiterative refinementtable-top manipulation tasks

Abstract

To make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the underlying plans, and executing actions in accordance with the devised plan through visual-motor control. We propose Compositional Foundation Models for Hierarchical Planning (HiP), a foundation model which leverages multiple expert foundation model trained on language, vision and action data individually jointly together to solve long-horizon tasks. We use a large language model to construct symbolic plans that are grounded in the environment through a large video diffusion model. Generated video plans are then grounded to visual-motor control, through an inverse dynamics model that infers actions from generated videos. To enable effective reasoning within this hierarchy, we enforce consistency between the models via iterative refinement. We illustrate the efficacy and adaptability of our approach in three different long-horizon table-top manipulation tasks.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Compositional Foundation Models for Hierarchical Planning | TensorX