论文信息 - Exact and Inexact Subsampled Newton Methods for Optimization

Exact and Inexact Subsampled Newton Methods for Optimization

The paper studies the solution of stochastic optimization problems in which approximations to the gradient and Hessian are obtained through subsampling. We first consider Newton-like methods that employ these approximations and discuss how to coordinate the accuracy in the gradient and Hessian to yield a superlinear rate of convergence in expectation. The second part of the paper analyzes an inexact Newton method that solves linear systems approximately using the conjugate gradient (CG) method, and that samples the Hessian and not the gradient (the gradient is assumed to be exact). We provide a complexity analysis for this method based on the properties of the CG iteration and the quality of the Hessian approximation, and compare it with a method that employs a stochastic gradient iteration instead of the CG method. We report preliminary numerical results that illustrate the performance of inexact subsampled Newton methods on machine learning applications based on logistic regression.

J. Nocedal | R. Byrd | Raghu Bollapragada

[1] R. Dembo,et al. INEXACT NEWTON METHODS , 1982 .

[2] Gene H. Golub,et al. Matrix computations , 1983 .

[3] Dimitri P. Bertsekas,et al. Nonlinear Programming , 1997 .

[4] Denis J. Dean,et al. Comparative accuracies of artificial neural networks and discriminant analysis in predicting forest cover types from cartographic variables , 1999 .

[5] Stephen J. Wright,et al. Numerical Optimization , 2018, Fundamental Statistical Inference.

[6] Steve R. Gunn,et al. Result Analysis of the NIPS 2003 Feature Selection Challenge , 2004, NIPS.

[7] Léon Bottou,et al. On-line learning for very large data sets , 2005 .

[8] A. Asuncion,et al. UCI Machine Learning Repository, University of California, Irvine, School of Information and Computer Sciences , 2007 .

[9] Léon Bottou,et al. The Tradeoffs of Large Scale Learning , 2007, NIPS.

[10] James Martens,et al. Deep learning via Hessian-free optimization , 2010, ICML.

[11] Stephen J. Wright,et al. Computational Methods for Sparse Solution of Linear Inverse Problems , 2010, Proceedings of the IEEE.