AADL: Anderson Accelerated Deep Learning
We propose a stable, distributed approach to perform AA that accelerates the convergence rate of stochastic first-order optimizers to train neural networks. Differently from previous works, we do not alter neither the scheme to perform AA nor the loss function minimized during the training. To improve robustness against stagnation, we customize general guidelines that suggest to relax the frequency of AA corrections by performing AA only at the end of an entire training epoch. To improve robustness of AA against the stochastic oscillations of first-order optimizers, we average the gradients computed on consecutive stochastic optimization updates. The improved regularity of the converging sequence and the reduced amplitude of stochastic oscillations across consecutive optimization steps allows AA to efficiently extrapolate an improved converging sequence, thereby overcoming limitations of existing approaches to perform AA on stochastic optimization.