论文信息 - Supervised Classification Using Sparse Fisher's LDA

Supervised Classification Using Sparse Fisher's LDA

It is well known that in a supervised classification setting when the number of features is smaller than the number of observations, Fisher's linear discriminant rule is asymptotically Bayes. However, there are numerous modern applications where classification is needed in the high-dimensional setting. Naive implementation of Fisher's rule in this case fails to provide good results because the sample covariance matrix is singular. Moreover, by constructing a classifier that relies on all features the interpretation of the results is challenging. Our goal is to provide robust classification that relies only on a small subset of important features and accounts for the underlying correlation structure. We apply a lasso-type penalty to the discriminant vector to ensure sparsity of the solution and use a shrinkage type estimator for the covariance matrix. The resulting optimization problem is solved using an iterative coordinate ascent algorithm. Furthermore, we analyze the effect of nonconvexity on the sparsity level of the solution and highlight the difference between the penalized and the constrained versions of the problem. The simulation results show that the proposed method performs favorably in comparison to alternatives. The method is used to classify leukemia patients based on DNA methylation features.

J. Booth | M. Wells | Irina Gaynanova

[1] Adam J. Rothman. Positive definite estimators of large covariance matrices , 2012 .

[2] J. Booth,et al. Integrative Model-based clustering of microarray methylation and expression data , 2012, 1210.0702.

[3] H. Zou,et al. A direct approach to sparse discriminant analysis in ultra-high dimensions , 2012 .

[4] Yang Feng,et al. A road to classification in high dimensional space: the regularized optimal affine discriminant , 2010, Journal of the Royal Statistical Society. Series B, Statistical methodology.

[5] R. Tibshirani,et al. Penalized classification using Fisher's linear discriminant , 2011, Journal of the Royal Statistical Society. Series B, Statistical methodology.

[6] Trevor J. Hastie,et al. Sparse Discriminant Analysis , 2011, Technometrics.

[7] T. Cai,et al. A Direct Estimation Approach to Sparse Linear Discriminant Analysis , 2011, 1107.3442.

[8] J. Shao,et al. Sparse linear discriminant analysis by thresholding for high dimensional data , 2011, 1105.3561.

[9] Guanghua Xiao,et al. Modeling Three-Dimensional Chromosome Structures Using Gene Expression Data , 2011, Journal of the American Statistical Association.

[10] Tiejun Tong,et al. Bias‐Corrected Diagonal Discriminant Rules for High‐Dimensional Classification , 2010, Biometrics.

[11] Fabien Campagne,et al. DNA methylation signatures identify biologically distinct subtypes in acute myeloid leukemia. , 2010, Cancer cell.