论文信息 - Evaluation of missing data imputation in longitudinal cohort studies in breast cancer survival

Evaluation of missing data imputation in longitudinal cohort studies in breast cancer survival

Missing values are common in medical datasets and may be amenable to data imputation when modelling a given data set or validating on an external cohort. This paper discusses model averaging over samples of the imputed distribution and extends this approach to generic non-linear modelling with the partial logistic artificial neural network (PLANN) regularised with automatic relevance determination (ARD). The study then applies the imputation to external validation, considering also predictions made for individual patients. A prognostic index is defined for the non-linear model and validation results show that four statistically significant risk groups identified at 95% level of confidence from the modelling data, from Christie Hospital (n = 931), retain good separation during external validation with data from the BC Cancer Agency (BCCA) (n = 4,083). A satisfactory discrimination and calibration performance was assessed with the time dependent C index (C td) and Hosmer-Lemeshow statistic, respectively, for both, training and validated model.

[1] Paulo J. G. Lisboa,et al. A Bayesian neural network approach for modelling censored data with an application to prognosis after surgery for breast cancer , 2003, Artif. Intell. Medicine.

[2] D. Cox. Regression Models and Life-Tables , 1972 .

[3] Karen A Gelmon,et al. Population-based validation of the prognostic model ADJUVANT! for early breast cancer. , 2005, Journal of clinical oncology : official journal of the American Society of Clinical Oncology.

[4] Douglas G Altman,et al. Developing a prognostic model in the presence of missing data: an ovarian cancer case study. , 2003, Journal of clinical epidemiology.

[5] H. Boshuizen,et al. Multiple imputation of missing blood pressure covariates in survival analysis. , 1999, Statistics in medicine.

[6] Ralph B. D'Agostino,et al. Evaluation of the Performance of Survival Analysis Models: Discrimination and Calibration Measures , 2003, Advances in Survival Analysis.

[7] J. Schafer. Multiple imputation: a primer , 1999, Statistical methods in medical research.

[8] Paulo J. G. Lisboa,et al. Double-blind evaluation and benchmarking of survival models in a multi-centre study , 2007, Comput. Biol. Medicine.

[9] Paulo J. G. Lisboa,et al. Stratification of Severity of Illness Indices: A Case Study for Breast Cancer Prognosis , 2008, KES.

[10] Paulo J. G. Lisboa,et al. A review of evidence of health benefit from artificial neural networks in medical intervention , 2002, Neural Networks.