Skip to main navigation Skip to search Skip to main content

Variance-reduced off-policy TDC learning: Non-asymptotic convergence analysis

  • Shaocong Ma
  • , Yi Zhou
  • , Shaofeng Zou
  • University of Utah

Research output: Contribution to journalConference articlepeer-review

12 Scopus citations

Abstract

Variance reduction techniques have been successfully applied to temporal-difference (TD) learning and help to improve the sample complexity in policy evaluation. However, the existing work applied variance reduction to either the less popular one time-scale TD algorithm or the two time-scale GTD algorithm but with a finite number of i.i.d. samples, and both algorithms apply to only the on-policy setting. In this work, we develop a variance reduction scheme for the two time-scale TDC algorithm in the off-policy setting and analyze its non-asymptotic convergence rate over both i.i.d. and Markovian samples. In the i.i.d. setting, our algorithm achieves a sample complexity O(e- 3 5 log e-1) that is lower than the state-of-the-art result O(e-1 log e-1). In the Markovian setting, our algorithm achieves the state-of-the-art sample complexity O(e-1 log e-1) that is near-optimal. Experiments demonstrate that the proposed variance-reduced TDC achieves a smaller asymptotic convergence error than both the conventional TDC and the variance-reduced TD.

Original languageEnglish
JournalAdvances in Neural Information Processing Systems
Volume2020-December
StatePublished - 2020
Event34th Conference on Neural Information Processing Systems, NeurIPS 2020 - Virtual, Online
Duration: Dec 6 2020Dec 12 2020

Fingerprint

Dive into the research topics of 'Variance-reduced off-policy TDC learning: Non-asymptotic convergence analysis'. Together they form a unique fingerprint.

Cite this