TensorX
返回文献探索

Paper · arXiv 2307.09668

Towards A Unified Agent with Foundation Models

Norman Di Palo, Arunkumar Byravan, Leonard Hasenclever, Markus Wulfmeier, Nicolas Heess, Martin Riedmiller

14 upvotesJuly 18, 2023arXiv 预印本
AI 摘要

A framework using language as a reasoning tool in RL agents improves exploration efficiency and data reuse, and enhances the acquisition of skills for novel tasks.

language modelsvision language modelsreinforcement learning (RL)efficient explorationexperience data reusescheduling skillslearning from observationssparse-reward environmentrobotic manipulationoffline datasetsskill reuseimitating human experts

Abstract

Language Models and Vision Language Models have recently demonstrated unprecedented capabilities in terms of understanding human intentions, reasoning, scene understanding, and planning-like behaviour, in text form, among many others. In this work, we investigate how to embed and leverage such abilities in Reinforcement Learning (RL) agents. We design a framework that uses language as the core reasoning tool, exploring how this enables an agent to tackle a series of fundamental RL challenges, such as efficient exploration, reusing experience data, scheduling skills, and learning from observations, which traditionally require separate, vertically designed algorithms. We test our method on a sparse-reward simulated robotic manipulation environment, where a robot needs to stack a set of objects. We demonstrate substantial performance improvements over baselines in exploration efficiency and ability to reuse data from offline datasets, and illustrate how to reuse learned skills to solve novel tasks or imitate videos of human experts.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号