TensorX
返回文献探索

Paper · arXiv 2505.16925

Risk-Averse Reinforcement Learning with Itakura-Saito Loss

Igor Udovichenko, Olivier Croissant, Anita Toleutaeva, Evgeny Burnaev, Alexander Korotin

26 upvotesMay 22, 2025arXiv 预印本
AI 摘要

Proposed Itakura-Saito divergence-based loss function enhances numerical stability in risk-averse reinforcement learning using exponential utility functions.

reinforcement learningrisk-averseutility theoryexponential utility functionBellman equationsnumerical instabilityItakura-Saito divergencestate-value functionsaction-value functions

Abstract

Risk-averse reinforcement learning finds application in various high-stakes fields. Unlike classical reinforcement learning, which aims to maximize expected returns, risk-averse agents choose policies that minimize risk, occasionally sacrificing expected value. These preferences can be framed through utility theory. We focus on the specific case of the exponential utility function, where we can derive the Bellman equations and employ various reinforcement learning algorithms with few modifications. However, these methods suffer from numerical instability due to the need for exponent computation throughout the process. To address this, we introduce a numerically stable and mathematically sound loss function based on the Itakura-Saito divergence for learning state-value and action-value functions. We evaluate our proposed loss function against established alternatives, both theoretically and empirically. In the experimental section, we explore multiple financial scenarios, some with known analytical solutions, and show that our loss function outperforms the alternatives.

北京市昌平区探索星信息技术及软件开发工作室

京ICP备2026059466号
Risk-Averse Reinforcement Learning with Itakura-Saito Loss | TensorX