Skip to main navigation Skip to search Skip to main content

Theoretical Study of Conflict-Avoidant Multi-Objective Reinforcement Learning

  • Yudan Wang
  • , Peiyao Xiao
  • , Hao Ban
  • , Kaiyi Ji
  • , Shaofeng Zou
  • Arizona State University
  • SUNY Buffalo

Research output: Contribution to journalArticlepeer-review

1 Scopus citations

Abstract

Multi-objective reinforcement learning (MORL) has shown great promise in many real-world applications. Existing MORL algorithms often aim to learn a policy that optimizes individual objective functions simultaneously with a given prior preference (or weights) on different objectives. However, these methods often suffer from the issue of gradient conflict such that the objectives with larger gradients dominate the update direction, resulting in a performance degeneration on other objectives. In this paper, we develop a novel dynamic weighting multi-objective actor-critic algorithm (MOAC) under two options of sub-procedures named as conflict-avoidant (CA) and faster convergence (FC) in objective weight updates. MOAC-CA aims to find a CA update direction that maximizes the minimum value improvement among objectives, and MOAC-FC targets at a much faster convergence rate. We provide a comprehensive finite-time convergence analysis for both algorithms. We show that MOAC-CA can find a ϵ+ϵapp-accurate Pareto stationary policy using O(ϵ-5) samples, while ensuring a small ϵ+√ϵapp -level CA distance (defined as the distance to the CA direction), where ϵapp is the function approximation error. The analysis also shows that MOAC-FC improves the sample complexity to O(ϵ-3), but with a constant-level CA distance. Our experiments on MT10 demonstrate the improved performance of our algorithms over existing MORL methods with fixed preference.

Original languageEnglish
Pages (from-to)7254-7269
Number of pages16
JournalIEEE Transactions on Information Theory
Volume71
Issue number9
DOIs
StatePublished - 2025

Keywords

  • Pareto stationary
  • actor-critic
  • dynamic weight
  • finite sample analysis
  • linear function approximation

Fingerprint

Dive into the research topics of 'Theoretical Study of Conflict-Avoidant Multi-Objective Reinforcement Learning'. Together they form a unique fingerprint.

Cite this