On momentum methods and acceleration in stochastic optimization
It is well known that momentum gradient methods (e.g., Polyak’s heavy ball, Nesterov’s acceleration) yield significant improvements over vanilla gradient descent in deterministic optimization (i.e., where we have access to exact gradient of the function…