On the Promise of the Stochastic Generalized Gauss-Newton Method for
Training DNNs
- URL: http://arxiv.org/abs/2006.02409v4
- Date: Tue, 9 Jun 2020 08:58:08 GMT
- Title: On the Promise of the Stochastic Generalized Gauss-Newton Method for
Training DNNs
- Authors: Matilde Gargiani, Andrea Zanelli, Moritz Diehl, Frank Hutter
- Abstract summary: We study a generalized Gauss-Newton method (SGN) for training DNNs.
SGN is a second-order optimization method, with efficient iterations, that we demonstrate to often require substantially fewer iterations than standard SGD to converge.
We show that SGN does not only substantially improve over SGD in terms of the number of iterations, but also in terms of runtime.
This is made possible by an efficient, easy-to-use and flexible implementation of SGN we propose in the Theano deep learning platform.
- Score: 37.96456928567548
- License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
- Abstract: Following early work on Hessian-free methods for deep learning, we study a
stochastic generalized Gauss-Newton method (SGN) for training DNNs. SGN is a
second-order optimization method, with efficient iterations, that we
demonstrate to often require substantially fewer iterations than standard SGD
to converge. As the name suggests, SGN uses a Gauss-Newton approximation for
the Hessian matrix, and, in order to compute an approximate search direction,
relies on the conjugate gradient method combined with forward and reverse
automatic differentiation. Despite the success of SGD and its first-order
variants, and despite Hessian-free methods based on the Gauss-Newton Hessian
approximation having been already theoretically proposed as practical methods
for training DNNs, we believe that SGN has a lot of undiscovered and yet not
fully displayed potential in big mini-batch scenarios. For this setting, we
demonstrate that SGN does not only substantially improve over SGD in terms of
the number of iterations, but also in terms of runtime. This is made possible
by an efficient, easy-to-use and flexible implementation of SGN we propose in
the Theano deep learning platform, which, unlike Tensorflow and Pytorch,
supports forward automatic differentiation. This enables researchers to further
study and improve this promising optimization technique and hopefully
reconsider stochastic second-order methods as competitive optimization
techniques for training DNNs; we also hope that the promise of SGN may lead to
forward automatic differentiation being added to Tensorflow or Pytorch. Our
results also show that in big mini-batch scenarios SGN is more robust than SGD
with respect to its hyperparameters (we never had to tune its step-size for our
benchmarks!), which eases the expensive process of hyperparameter tuning that
is instead crucial for the performance of first-order methods.
Related papers
- Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor Generalization [52.16435732772263]
Second-order optimization has been shown to accelerate the training of deep neural networks in many applications.
However, generalization properties of second-order methods are still being debated.
We show for the first time that exact Gauss-Newton (GN) updates take on a tractable form in a class of deep architectures.
arXiv Detail & Related papers (2024-11-12T17:58:40Z) - Incremental Gauss-Newton Descent for Machine Learning [0.0]
We present a modification of the Gradient Descent algorithm exploiting approximate second-order information based on the Gauss-Newton approach.
The new method, which we call Incremental Gauss-Newton Descent (IGND), has essentially the same computational burden as standard SGD.
IGND can significantly outperform SGD while performing at least as well as SGD in the worst case.
arXiv Detail & Related papers (2024-08-10T13:52:40Z) - Exact Gauss-Newton Optimization for Training Deep Neural Networks [0.0]
We present EGN, a second-order optimization algorithm that combines the generalized Gauss-Newton (GN) Hessian approximation with low-rank linear algebra to compute the descent direction.
We show how improvements such as line search, adaptive regularization, and momentum can be seamlessly added to EGN to further accelerate the algorithm.
arXiv Detail & Related papers (2024-05-23T10:21:05Z) - Gauss-Newton Temporal Difference Learning with Nonlinear Function Approximation [11.925232472331494]
A Gauss-Newton Temporal Difference (GNTD) learning method is proposed to solve the Q-learning problem with nonlinear function approximation.
In each iteration, our method takes one Gauss-Newton (GN) step to optimize a variant of Mean-Squared Bellman Error (MSBE)
We validate our method via extensive experiments in several RL benchmarks, where GNTD exhibits both higher rewards and faster convergence than TD-type methods.
arXiv Detail & Related papers (2023-02-25T14:14:01Z) - Positive-Negative Momentum: Manipulating Stochastic Gradient Noise to
Improve Generalization [89.7882166459412]
gradient noise (SGN) acts as implicit regularization for deep learning.
Some works attempted to artificially simulate SGN by injecting random noise to improve deep learning.
For simulating SGN at low computational costs and without changing the learning rate or batch size, we propose the Positive-Negative Momentum (PNM) approach.
arXiv Detail & Related papers (2021-03-31T16:08:06Z) - Direction Matters: On the Implicit Bias of Stochastic Gradient Descent
with Moderate Learning Rate [105.62979485062756]
This paper attempts to characterize the particular regularization effect of SGD in the moderate learning rate regime.
We show that SGD converges along the large eigenvalue directions of the data matrix, while GD goes after the small eigenvalue directions.
arXiv Detail & Related papers (2020-11-04T21:07:52Z) - Improving predictions of Bayesian neural nets via local linearization [79.21517734364093]
We argue that the Gauss-Newton approximation should be understood as a local linearization of the underlying Bayesian neural network (BNN)
Because we use this linearized model for posterior inference, we should also predict using this modified model instead of the original one.
We refer to this modified predictive as "GLM predictive" and show that it effectively resolves common underfitting problems of the Laplace approximation.
arXiv Detail & Related papers (2020-08-19T12:35:55Z) - Deep Neural Network Learning with Second-Order Optimizers -- a Practical
Study with a Stochastic Quasi-Gauss-Newton Method [0.0]
We introduce and study a second-order quasi-Gauss-Newton (SQGN) optimization method that combines ideas from quasi-Newton methods, Gauss-Newton methods, and variance reduction to address this problem.
We discuss the implementation of SQGN with benchmark, and we compare its convergence and computational performance to selected first-order methods.
arXiv Detail & Related papers (2020-04-06T23:41:41Z) - Learning to Optimize Non-Rigid Tracking [54.94145312763044]
We employ learnable optimizations to improve robustness and speed up solver convergence.
First, we upgrade the tracking objective by integrating an alignment data term on deep features which are learned end-to-end through CNN.
Second, we bridge the gap between the preconditioning technique and learning method by introducing a ConditionNet which is trained to generate a preconditioner.
arXiv Detail & Related papers (2020-03-27T04:40:57Z)
This list is automatically generated from the titles and abstracts of the papers in this site.
This site does not guarantee the quality of this site (including all information) and is not responsible for any consequences.