Skip to main navigation Skip to search Skip to main content

Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization

  • Yufeng Yang
  • , Erin E. Tripp
  • , Yifan Sun
  • , Shaofeng Zou
  • , Yi Zhou
  • Texas A&M University
  • Hamilton College
  • Stony Brook University

Research output: Contribution to journalArticlepeer-review

Abstract

Recent studies have shown that many nonconvex machine learning problems satisfy a generalized-smooth condition that extends beyond traditional smooth nonconvex optimization. However, the existing algorithms are not fully adapted to such generalized-smooth nonconvex geometry and encounter significant technical limitations on their convergence analysis. In this work, we first analyze the convergence of adaptively normalized gradient descent under function geometries characterized by generalized-smoothness and generalized PŁ condition, revealing the advantage of adaptive gradient normalization. Our results provide theoretical insights into adaptive normalization across various scenarios. For stochastic generalized-smooth nonconvex optimization, we propose Independent-Adaptively Normalized Stochastic Gradient Descent algorithm, which leverages adaptive gradient normalization, independent sampling, and gradient clipping to achieve an O(ϵ−4) sample complexity under relaxed noise assumptions. Experiments1 on large-scale nonconvex generalized-smooth problems demonstrate the fast convergence of our algorithm.

Original languageEnglish
JournalTransactions on Machine Learning Research
Volume2025-July
StatePublished - 2025

Fingerprint

Dive into the research topics of 'Adaptive Gradient Normalization and Independent Sampling for (Stochastic) Generalized-Smooth Optimization'. Together they form a unique fingerprint.

Cite this