论文信息 - Scaling Large Learning Problems with Hard Parallel Mixtures

Scaling Large Learning Problems with Hard Parallel Mixtures

A challenge for statistical learning is to deal with large data sets, e.g. in data mining. Popular learning algorithms such as Support Vector Machines have training time at least quadratic in the number of examples: they are hopeless to solve problems with a million examples. We propose a "hard parallelizable mixture" methodology which yields significantly reduced training time through modularization and parallelization: the training data is iteratively partitioned by a "gater" model in such a way that it becomes easy to learn an "expert" model separately in each region of the partition. Ap robabilistic extension and the use of a set of generative models allows representing the gater so that all pieces of the model are locally trained. For SVMs, time complexity appears empirically to locally grow linearly with the number of examples, while generalization performance can be enhanced. For the probabilistic version of the algorithm, the iterative algorithm provably goes down in a cost function that is an upper bound on the negative log-likelihood.

Samy Bengio | Yoshua Bengio | Ronan Collobert

[1] Geoffrey E. Hinton,et al. Adaptive Mixtures of Local Experts , 1991, Neural Computation.

[2] John Platt,et al. Probabilistic Outputs for Support vector Machines and Comparisons to Regularized Likelihood Methods , 1999 .

[3] James T. Kwok. Support vector mixture for classification and regression problems , 1998, ICPR.

[4] Volker Tresp,et al. A Bayesian Committee Machine , 2000, Neural Computation.

[5] Mitsuo Kawato,et al. MOSAIC Model for Sensorimotor Learning and Control , 2001, Neural Computation.

[6] Samy Bengio,et al. SVMTorch: Support Vector Machines for Large-Scale Regression Problems , 2001, J. Mach. Learn. Res..

[7] Vladimir N. Vapnik,et al. The Nature of Statistical Learning Theory , 2000, Statistics for Engineering and Information Science.

[8] Christian Pellegrini,et al. Local experts combination through density decomposition , 1999, AISTATS.

[9] Ronald A. Cole,et al. New telephone speech corpora at CSLU , 1995, EUROSPEECH.

[10] Carl E. Rasmussen,et al. In Advances in Neural Information Processing Systems , 2011 .