论文信息 - Super Learning: An Application to the Prediction of HIV-1 Drug Resistance

Super Learning: An Application to the Prediction of HIV-1 Drug Resistance

Many alternative data-adaptive algorithms can be used to learn a predictor based on observed data. Examples of such learners include decision trees, neural networks, support vector regression, least angle regression, logic regression, and the Deletion/Substitution/Addition algorithm. The optimal learner for prediction will vary depending on the underlying data-generating distribution. In this article we introduce the "super learner", a prediction algorithm that applies any set of candidate learners and uses cross-validation to select between them. Theory shows that asymptotically the super learner performs essentially as well as or better than any of the candidate learners. In this article we present the theory behind the super learner, and illustrate its performance using simulations. We further apply the super learner to a data example, in which we predict the phenotypic antiretroviral susceptibility of HIV based on viral genotype. Specifically, we apply the super learner to predict susceptibility to a specific protease inhibitor, nelfinavir, using a set of database-derived non-polymorphic treatment-selected mutations.

M. J. van der Laan | E. Polley | S. Sinisi | M. Petersen | Soo-Yon Rhee

[1] R. Shafer,et al. Genotypic predictors of human immunodeficiency virus type 1 drug resistance , 2006, Proceedings of the National Academy of Sciences.

[2] M. J. Laan. Statistical Inference for Variable Importance , 2006 .

[3] Tommy F. Liu,et al. HIV-1 Protease and reverse-transcriptase mutations: correlations with antiretroviral therapy in subtype B isolates and implications for drug-resistance surveillance. , 2005, The Journal of infectious diseases.

[4] Mark J van der Laan,et al. Deletion/Substitution/Addition Algorithm in Learning with Applications in Genomics , 2004, Statistical applications in genetics and molecular biology.

[5] Arthur E. Hoerl,et al. Ridge Regression: Biased Estimation for Nonorthogonal Problems , 2000, Technometrics.

[6] A. E. Hoerl,et al. Ridge Regression: Applications to Nonorthogonal Problems , 1970 .

[7] M. J. Laan,et al. Application of a Variable Importance Measure Method to HIV-1 Sequence Data , 2005 .

[8] R. Tibshirani,et al. Least angle regression , 2004, math/0406456.