Skip to main navigation Skip to search Skip to main content

Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis

  • Ziyi Chen
  • , Yi Zhou
  • , Rong Rong Chen
  • , Shaofeng Zou
  • University of Utah

Research output: Contribution to journalConference articlepeer-review

28 Scopus citations

Abstract

Actor-critic (AC) algorithms have been widely used in decentralized multi-agent systems to learn the optimal joint control policy. However, existing decentralized AC algorithms either need to share agents' sensitive information or lack communication-efficiency. In this work, we develop decentralized AC and natural AC (NAC) algorithms that avoid sharing agents' local information and are sample and communication-efficient. In both algorithms, agents share only noisy rewards and use mini-batch local policy gradient updates to ensure high sample and communication efficiency. Particularly for decentralized NAC, we develop a decentralized Markovian SGD algorithm with an adaptive mini-batch size to efficiently compute the natural policy gradient. Under Markovian sampling and linear function approximation, we prove that the proposed decentralized AC and NAC algorithms achieve the state-of-the-art sample complexities O(ϵ−2 ln ϵ−1) and O(ϵ−3 ln ϵ−1), respectively, and achieve an improved communication complexity O(ϵ−1 ln ϵ−1). Numerical experiments demonstrate that the proposed algorithms achieve lower sample and communication complexities than the existing decentralized AC algorithms.

Original languageEnglish
Pages (from-to)3794-3834
Number of pages41
JournalProceedings of Machine Learning Research
Volume162
StatePublished - 2022
Event39th International Conference on Machine Learning, ICML 2022 - Baltimore, United States
Duration: Jul 17 2022Jul 23 2022

Fingerprint

Dive into the research topics of 'Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time Analysis'. Together they form a unique fingerprint.

Cite this