TensorX
返回文献探索

Paper · arXiv 2501.14176

RL + Transformer = A General-Purpose Problem Solver

Micah Rentschler, Jesse Roberts

28 upvotesJanuary 24, 2025arXiv 预印本
AI 摘要

A pre-trained transformer fine-tuned with reinforcement learning develops the ability to solve unseen problems with sample efficiency and adaptability, demonstrating robust meta-learning.

transformerreinforcement learningIn-Context Reinforcement Learning (ICRL)meta-learnersample efficiencyout-of-distributionrobustnessbehavior stitchingnon-stationary environments

Abstract

What if artificial intelligence could not only solve problems for which it was trained but also learn to teach itself to solve new problems (i.e., meta-learn)? In this study, we demonstrate that a pre-trained transformer fine-tuned with reinforcement learning over multiple episodes develops the ability to solve problems that it has never encountered before - an emergent ability called In-Context Reinforcement Learning (ICRL). This powerful meta-learner not only excels in solving unseen in-distribution environments with remarkable sample efficiency, but also shows strong performance in out-of-distribution environments. In addition, we show that it exhibits robustness to the quality of its training data, seamlessly stitches together behaviors from its context, and adapts to non-stationary environments. These behaviors demonstrate that an RL-trained transformer can iteratively improve upon its own solutions, making it an excellent general-purpose problem solver.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
RL + Transformer = A General-Purpose Problem Solver | TensorX