Skip to main navigation Skip to search Skip to main content

Robust Average-Reward Reinforcement Learning

  • Yue Wang
  • , Alvaro Velasquez
  • , George Atia
  • , Ashley Prater-Bennette
  • , Shaofeng Zou
  • University of Central Florida
  • University of Colorado Boulder
  • Air Force Research Laboratory

Research output: Contribution to journalArticlepeer-review

7 Scopus citations

Abstract

Robust Markov decision processes (MDPs) aim to find a policy that optimizes the worst-case performance over an uncertainty set of MDPs. Existing studies mostly have focused on the robust MDPs under the discounted reward criterion, leaving the ones under the average-reward criterion largely unexplored. In this paper, we develop the first comprehensive and systematic study of robust average-reward MDPs, where the goal is to optimize the long-term average performance under the worst case. Our contributions are four-folds: (1) we prove the uniform convergence of the robust discounted value function to the robust average-reward function as the discount factor γ goes to 1; (2) we derive the robust average-reward Bellman equation, characterize the structure of its solution set, and prove the equivalence between solving the robust Bellman equation and finding the optimal robust policy; (3) we design robust dynamic programming algorithms, and theoretically characterize their convergence to the optimal policy; and (4) we design two model-free algorithms unitizing the multi-level Monte-Carlo approach, and prove their asymptotic convergence.

Original languageEnglish
Pages (from-to)719-803
Number of pages85
JournalJournal of Artificial Intelligence Research
Volume80
DOIs
StatePublished - 2024

Fingerprint

Dive into the research topics of 'Robust Average-Reward Reinforcement Learning'. Together they form a unique fingerprint.

Cite this