Skip to main navigation Skip to search Skip to main content

Policy Gradient Method For Robust Reinforcement Learning

  • Yue Wang
  • , Shaofeng Zou
  • SUNY Buffalo

Research output: Contribution to journalConference articlepeer-review

60 Scopus citations

Abstract

This paper develops the first policy gradient method with global optimality guarantee and complexity analysis for robust reinforcement learning under model mismatch. Robust reinforcement learning is to learn a policy robust to model mismatch between simulator and real environment. We first develop the robust policy (sub-)gradient, which is applicable for any differentiable parametric policy class. We show that the proposed robust policy gradient method converges to the global optimum asymptotically under direct policy parameterization. We further develop a smoothed robust policy gradient method, and show that to achieve an ϵ-global optimum, the complexity is O(ϵ−3). We then extend our methodology to the general model-free setting, and design the robust actor-critic method with differentiable parametric policy class and value function. We further characterize its asymptotic convergence and sample complexity under the tabular setting. Finally, we provide simulation results to demonstrate the robustness of our methods.

Original languageEnglish
Pages (from-to)23484-23526
Number of pages43
JournalProceedings of Machine Learning Research
Volume162
StatePublished - 2022
Event39th International Conference on Machine Learning, ICML 2022 - Baltimore, United States
Duration: Jul 17 2022Jul 23 2022

Fingerprint

Dive into the research topics of 'Policy Gradient Method For Robust Reinforcement Learning'. Together they form a unique fingerprint.

Cite this