TensorX
返回文献探索

Paper · arXiv 2310.10625

Video Language Planning

Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B. Tenenbaum, Leslie Kaelbling, Andy Zeng, Jonathan Tompson

11 upvotesOctober 16, 2023arXiv 预印本
AI 摘要

Video language planning leverages vision-language and text-to-video models to generate detailed long-horizon video plans, improving task success rates across multiple robotics domains.

video language planningtree search procedurevision-language modelsvalue functionstext-to-video modelsdynamics modelsgoal-conditioned policieslong-horizon tasksmulti-object rearrangementmulti-camera bi-arm dexterous manipulationrobot actions

Abstract

We are interested in enabling visual planning for complex long-horizon tasks in the space of generated videos and language, leveraging recent advances in large generative models pretrained on Internet-scale data. To this end, we present video language planning (VLP), an algorithm that consists of a tree search procedure, where we train (i) vision-language models to serve as both policies and value functions, and (ii) text-to-video models as dynamics models. VLP takes as input a long-horizon task instruction and current image observation, and outputs a long video plan that provides detailed multimodal (video and language) specifications that describe how to complete the final task. VLP scales with increasing computation budget where more computation time results in improved video plans, and is able to synthesize long-horizon video plans across different robotics domains: from multi-object rearrangement, to multi-camera bi-arm dexterous manipulation. Generated video plans can be translated into real robot actions via goal-conditioned policies, conditioned on each intermediate frame of the generated video. Experiments show that VLP substantially improves long-horizon task success rates compared to prior methods on both simulated and real robots (across 3 hardware platforms).

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Video Language Planning | TensorX