论文信息 - Flexible Data Trimming Improves Performance of Global Machine Learning Methods in Omics-Based Personalized Oncology

Flexible Data Trimming Improves Performance of Global Machine Learning Methods in Omics-Based Personalized Oncology

(1) Background: Machine learning (ML) methods are rarely used for an omics-based prescription of cancer drugs, due to shortage of case histories with clinical outcome supplemented by high-throughput molecular data. This causes overtraining and high vulnerability of most ML methods. Recently, we proposed a hybrid global-local approach to ML termed floating window projective separator (FloWPS) that avoids extrapolation in the feature space. Its core property is data trimming, i.e., sample-specific removal of irrelevant features. (2) Methods: Here, we applied FloWPS to seven popular ML methods, including linear SVM, k nearest neighbors (kNN), random forest (RF), Tikhonov (ridge) regression (RR), binomial naïve Bayes (BNB), adaptive boosting (ADA) and multi-layer perceptron (MLP). (3) Results: We performed computational experiments for 21 high throughput gene expression datasets (41–235 samples per dataset) totally representing 1778 cancer patients with known responses on chemotherapy treatments. FloWPS essentially improved the classifier quality for all global ML methods (SVM, RF, BNB, ADA, MLP), where the area under the receiver-operator curve (ROC AUC) for the treatment response classifiers increased from 0.61–0.88 range to 0.70–0.94. We tested FloWPS-empowered methods for overtraining by interrogating the importance of different features for different ML methods in the same model datasets. (4) Conclusions: We showed that FloWPS increases the correlation of feature importance between the different ML methods, which indicates its robustness to overtraining. For all the datasets tested, the best performance of FloWPS data trimming was observed for the BNB method, which can be valuable for further building of ML classifiers in personalized oncology.

[1] Bhupinder Bhullar,et al. Molecular pathway activation features linked with transition from normal skin to primary and metastatic melanomas in human , 2015, Oncotarget.

[2] Nikolay Borisov,et al. Individual Drug Treatment Prediction in Oncology Based on Machine Learning Using Cell Culture Gene Expression Data , 2017, ICCBB.

[3] Gastone Castellani,et al. The genetic and genomic background of multiple myeloma patients achieving complete response after induction therapy with bortezomib, thalidomide and dexamethasone (VTD) , 2015, Oncotarget.

[4] Anthony Boral,et al. Gene expression profiling and correlation with outcome in clinical trials of the proteasome inhibitor bortezomib. , 2006, Blood.

[5] Christopher M. Bishop,et al. Pattern Recognition and Machine Learning (Information Science and Statistics) , 2006 .

[6] M. Dowsett,et al. Accurate Prediction and Validation of Response to Endocrine Therapy in Breast Cancer. , 2015, Journal of clinical oncology : official journal of the American Society of Clinical Oncology.

[7] J. S. Cramer. The Origins of Logistic Regression , 2002 .

[8] Nicolas Borisov,et al. New Paradigm of Machine Learning (ML) in Personalized Oncology: Data Trimming for Squeezing More Biomarkers From Clinical Datasets , 2019, Front. Oncol..

[9] Nicolas Borisov,et al. Shambhala: a platform-agnostic data harmonizer for gene expression data , 2019, BMC Bioinformatics.

[10] Nikolay M. Borisov,et al. Pathway Based Analysis of Mutation Data Is Efficient for Scoring Target Cancer Drugs , 2019, Front. Pharmacol..

[11] G. Molenberghs,et al. Type I and Type II Error Under Random‐Effects Misspecification in Generalized Linear Mixed Models , 2007, Biometrics.

[12] Hae-Young Kim. Statistical notes for clinical researchers: Type I and type II errors in statistical decision , 2015, Restorative dentistry & endodontics.

[13] Yuan Qi,et al. Cell Line Derived Multi-Gene Predictor of Pathologic Response to Neoadjuvant Chemotherapy in Breast Cancer: A Validation Study on US Oncology 02-103 Clinical Trial , 2012, BMC Medical Genomics.

[14] J. Jakobsen,et al. Trial Sequential Analysis in systematic reviews with meta-analysis , 2017, BMC Medical Research Methodology.

[15] Amir Samii,et al. Molecular pathway activation - New type of biomarkers for tumor morphology and personalized selection of target drugs. , 2018, Seminars in cancer biology.

[16] Federico Girosi,et al. An improved training algorithm for support vector machines , 1997, Neural Networks for Signal Processing VII. Proceedings of the 1997 IEEE Signal Processing Society Workshop.

[17] John P A Ioannidis,et al. Optimal type I and type II error pairs when the available sample size is fixed. , 2013, Journal of clinical epidemiology.

[18] Mary Goldman,et al. The UCSC Cancer Genomics Browser: update 2015 , 2014, Nucleic Acids Res..

[19] M. Sorokin,et al. RNA sequencing for research and diagnostics in clinical oncology. , 2020, Seminars in cancer biology.

[20] T. Reynoldson,et al. Evaluating the Type II error rate in a sediment toxicity classification using the Reference Condition Approach. , 2011, Aquatic toxicology.

[21] S. Stigler,et al. The History of Statistics: The Measurement of Uncertainty before 1900 , 1986 .

[22] Nicolas Borisov,et al. High-Throughput Mutation Data Now Complement Transcriptomic Profiling: Advances in Molecular Pathway Activation Analysis Approach in Cancer Biology , 2019, Cancer informatics.

[23] Yuan Qi,et al. Gene pathways associated with prognosis and chemotherapy sensitivity in molecular subtypes of breast cancer. , 2011, Journal of the National Cancer Institute.

[24] Rieko Arimoto,et al. Development of CYP3A4 Inhibition Models: Comparisons of Machine-Learning Techniques and Molecular Descriptors , 2005, Journal of biomolecular screening.

[25] Nicolas Borisov,et al. A method for predicting target drug efficiency in cancer based on the analysis of signaling pathway activation , 2015, Oncotarget.

[26] M. Hazinski,et al. Guidelines based on fear of type II (false-negative) errors. Why we dropped the pulse check for lay rescuers. , 2000, Resuscitation.

[27] Yi Li,et al. Gene Expression Profile Alone Is Inadequate In Predicting Complete Response In Multiple Myeloma , 2014, Leukemia.

[28] David C. Atkins,et al. Identification of Molecular Predictors of Response in a Study of Tipifarnib Treatment in Relapsed and Refractory Acute Myelogenous Leukemia , 2007, Clinical Cancer Research.