Abstract
Variance-reduced algorithms, although achieve great theoretical performance, can run slowly in practice due to the periodic gradient estimation with a large batch of data. Batch-size adaptation thus arises as a promising approach to acceler ate such algorithms. However, existing schemes either apply prescribed batch-size adaption rule or exploit the information along optimization path via additional backtracking and condition verification steps. In this paper, we propose a novel scheme, which eliminates backtracking line search but still exploits the information along op timization path by adapting the batch size via his tory stochastic gradients. We further theoretically show that such a scheme substantially reduces the overall complexity for popular variance-reduced algorithms SVRG and SARAH/SPIDER for both conventional nonconvex optimization and rein forcement learning problems. To this end, we develop a new convergence analysis framework to handle the dependence of the batch size on his tory stochastic gradients. Extensive experiments validate the effectiveness of the proposed batch size adaptation scheme.
| Original language | English |
|---|---|
| Journal | Proceedings of Machine Learning Research |
| Volume | 119 |
| State | Published - 2020 |
| Event | 37th International Conference on Machine Learning, ICML 2020 - Virtual, Online Duration: Jul 13 2020 → Jul 18 2020 |
Fingerprint
Dive into the research topics of 'History-Gradient Aided Batch Size Adaptation for Variance Reduced Algorithms'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver