Losing momentum in continuous-time stochastic optimisation

Autor:	Jin, Kexin, Latz, Jonas, Liu, Chenguang, Scagliotti, Alessandro
Rok vydání:	2022
Předmět:	Mathematics - Optimization and Control Computer Science - Machine Learning Mathematics - Numerical Analysis 90C15 37N40 37H30 65C40 68T07 68W20
Druh dokumentu:	Working Paper
Popis:	The training of modern machine learning models often consists in solving high-dimensional non-convex optimisation problems that are subject to large-scale data. In this context, momentum-based stochastic optimisation algorithms have become particularly widespread. The stochasticity arises from data subsampling which reduces computational cost. Both, momentum and stochasticity help the algorithm to converge globally. In this work, we propose and analyse a continuous-time model for stochastic gradient descent with momentum. This model is a piecewise-deterministic Markov process that represents the optimiser by an underdamped dynamical system and the data subsampling through a stochastic switching. We investigate longtime limits, the subsampling-to-no-subsampling limit, and the momentum-to-no-momentum limit. We are particularly interested in the case of reducing the momentum over time. Under convexity assumptions, we show convergence of our dynamical system to the global minimiser when reducing momentum over time and letting the subsampling rate go to infinity. We then propose a stable, symplectic discretisation scheme to construct an algorithm from our continuous-time dynamical system. In experiments, we study our scheme in convex and non-convex test problems. Additionally, we train a convolutional neural network in an image classification problem. Our algorithm {attains} competitive results compared to stochastic gradient descent with momentum.
Databáze:	arXiv
Externí odkaz:	http://arxiv.org/abs/2209.03705 Zobrazit plný text záznamu View this record from Arxiv