论文信息 - Learning the Learning Rate for Prediction with Expert Advice

Learning the Learning Rate for Prediction with Expert Advice

Most standard algorithms for prediction with expert advice depend on a parameter called the learning rate. This learning rate needs to be large enough to fit the data well, but small enough to prevent overfitting. For the exponential weights algorithm, a sequence of prior work has established theoretical guarantees for higher and higher data-dependent tunings of the learning rate, which allow for increasingly aggressive learning. But in practice such theoretical tunings often still perform worse (as measured by their regret) than ad hoc tuning with an even higher learning rate. To close the gap between theory and practice we introduce an approach to learn the learning rate. Up to a factor that is at most (poly)logarithmic in the number of experts and the inverse of the learning rate, our method performs as well as if we would know the empirically best learning rate from a large range that includes both conservative small values and values that are much higher than those for which formal guarantees were previously available. Our method employs a grid of learning rates, yet runs in linear time regardless of the size of the grid.

Wouter M. Koolen | Peter Grünwald | Tim van Erven | P. Grünwald | T. Erven

[1] Peter Grünwald,et al. The Safe Bayesian - Learning the Learning Rate via the Mixability Gap , 2012, ALT.

[2] V. Vovk. Competitive On‐line Statistics , 2001 .

[3] Yoav Freund,et al. A decision-theoretic generalization of on-line learning and an application to boosting , 1995, EuroCOLT.

[4] Vladimir Vovk,et al. A game of prediction with expert advice , 1995, COLT '95.

[5] Claudio Gentile,et al. Adaptive and Self-Confident On-Line Learning Algorithms , 2000, J. Comput. Syst. Sci..

[6] Wouter M. Koolen,et al. Follow the leader if you can, hedge if you must , 2013, J. Mach. Learn. Res..

[7] Manfred K. Warmuth,et al. The Weighted Majority Algorithm , 1994, Inf. Comput..

[8] Gilles Stoltz,et al. Forecasting electricity consumption by aggregating specialized experts A review of the sequential aggregation of specialized experts, with an application to Slovakian and French country-wide one-day-ahead (half-)hourly predictions , 2012 .

[9] Thomas M. Cover,et al. Elements of Information Theory , 2005 .

[10] Gábor Lugosi,et al. Prediction, learning, and games , 2006 .

[11] Yishay Mansour,et al. Improved second-order bounds for prediction with expert advice , 2006, Machine Learning.

[12] Wouter M. Koolen,et al. Adaptive Hedge , 2011, NIPS.

[13] Gilles Stoltz,et al. Forecasting electricity consumption by aggregating specialized experts , 2012, Machine Learning.