TensorX
返回文献探索

Paper · arXiv 2501.16142

Towards General-Purpose Model-Free Reinforcement Learning

Scott Fujimoto, Pierluca D'Oro, Amy Zhang, Yuandong Tian, Michael Rabbat

31 upvotesJanuary 27, 2025arXiv 预印本
AI 摘要

The paper presents MR.Q, a model-free deep RL algorithm that uses model-based representations to improve performance across diverse benchmarks without incurring high computational costs.

reinforcement learningmodel-based RLvalue functiontask objectivesmodel-free RLMR.Q

Abstract

Reinforcement learning (RL) promises a framework for near-universal problem-solving. In practice however, RL algorithms are often tailored to specific benchmarks, relying on carefully tuned hyperparameters and algorithmic choices. Recently, powerful model-based RL methods have shown impressive general results across benchmarks but come at the cost of increased complexity and slow run times, limiting their broader applicability. In this paper, we attempt to find a unifying model-free deep RL algorithm that can address a diverse class of domains and problem settings. To achieve this, we leverage model-based representations that approximately linearize the value function, taking advantage of the denser task objectives used by model-based RL while avoiding the costs associated with planning or simulated trajectories. We evaluate our algorithm, MR.Q, on a variety of common RL benchmarks with a single set of hyperparameters and show a competitive performance against domain-specific and general baselines, providing a concrete step towards building general-purpose model-free deep RL algorithms.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号