Joint adaptive mean-variance regularization and variance stabilization of high dimensional data

The paper addresses a common problem in the analysis of high-dimensional high-throughput "omics" data, which is parameter estimation across multiple variables in a set of data where the number of variables is much larger than the sample size. Among the problems posed by this type of data are that variable-specific estimators of variances are not reliable and variable-wise tests statistics have low power, both due to a lack of degrees of freedom. In addition, it has been observed in this type of data that the variance increases as a function of the mean. We introduce a non-parametric adaptive regularization procedure that is innovative in that : (i) it employs a novel "similarity statistic"-based clustering technique to generate local-pooled or regularized shrinkage estimators of population parameters, (ii) the regularization is done jointly on population moments, benefiting from C. Stein's result on inadmissibility, which implies that usual sample variance estimator is improved by a shrinkage estimator using information contained in the sample mean. From these joint regularized shrinkage estimators, we derived regularized t-like statistics and show in simulation studies that they offer more statistical power in hypothesis testing than their standard sample counterparts, or regular common value-shrinkage estimators, or when the information contained in the sample mean is simply ignored. Finally, we show that these estimators feature interesting properties of variance stabilization and normalization that can be used for preprocessing high-dimensional multivariate data. The method is available as an R package, called 'MVR' ('Mean-Variance Regularization'), downloadable from the CRAN website.

[1]  W Y Zhang,et al.  Discussion on `Sure independence screening for ultra-high dimensional feature space' by Fan, J and Lv, J. , 2008 .

[2]  C. Stein Inadmissibility of the Usual Estimator for the Mean of a Multivariate Normal Distribution , 1956 .

[3]  Martin Vingron,et al.  Variance stabilization applied to microarray data calibration and to the quantification of differential expression , 2002, ISMB.

[4]  L. L. Cam Mathematical Statistics and Probability - Stochastic Processes. , 1976 .

[5]  Y. Benjamini,et al.  THE CONTROL OF THE FALSE DISCOVERY RATE IN MULTIPLE TESTING UNDER DEPENDENCY , 2001 .

[6]  C M Kendziorski,et al.  On parametric empirical Bayes methods for comparing multiple groups using replicated gene expression profiles , 2003, Statistics in medicine.

[7]  S. Knudsen,et al.  A new non-linear normalization method for reducing variability in DNA microarray experiments , 2002, Genome Biology.

[8]  Wing Hung Wong,et al.  TileMap: create chromosomal map of tiling array hybridizations , 2005, Bioinform..

[9]  Raymond J Carroll,et al.  Variance estimation in the analysis of microarray data , 2009, Journal of the Royal Statistical Society. Series B, Statistical methodology.

[10]  Korbinian Strimmer,et al.  Modeling gene expression measurement error: a quasi-likelihood approach , 2003, BMC Bioinformatics.

[11]  Michael L. Bittner,et al.  Ratio statistics of gene expression levels and applications to microarray data analysis , 2002, Bioinform..

[12]  B. Reinhard,et al.  Quantification of differential ErbB1 and ErbB2 cell surface expression and spatial nanoclustering through plasmon coupling. , 2012, Nano letters.

[13]  Robert Tibshirani,et al.  The Elements of Statistical Learning: Data Mining, Inference, and Prediction, 2nd Edition , 2001, Springer Series in Statistics.

[14]  John D. Storey,et al.  Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach , 2004 .

[15]  C. Stein Inadmissibility of the usual estimator for the variance of a normal distribution with unknown mean , 1964 .

[16]  L. Wasserman,et al.  Operating characteristics and extensions of the false discovery rate procedure , 2002 .

[17]  R. Tibshirani,et al.  On testing the significance of sets of genes , 2006, math/0610667.

[18]  Y. Benjamini,et al.  Controlling the false discovery rate: a practical and powerful approach to multiple testing , 1995 .

[19]  Ya'acov Ritov,et al.  Discussion: The Dantzig selector: statistical estimation when $p$ is much larger than $n$ , 2008, 0803.3130.

[20]  Chris A Glasbey,et al.  A comparison of parametric and nonparametric methods for normalising cDNA microarray data. , 2007, Biometrical journal. Biometrische Zeitschrift.

[21]  W. Cleveland Robust Locally Weighted Regression and Smoothing Scatterplots , 1979 .

[22]  R. Tibshirani,et al.  Significance analysis of microarrays applied to the ionizing radiation response , 2001, Proceedings of the National Academy of Sciences of the United States of America.

[23]  Hemant Ishwaran,et al.  CART variance stabilization and regularization for high-throughput genomic data , 2006, Bioinform..

[24]  J. S. Rao,et al.  Detecting Differentially Expressed Genes in Microarrays Using Bayesian Model Selection , 2003 .

[25]  Richard Simon,et al.  A random variance model for detection of differential gene expression in small microarray experiments , 2003, Bioinform..

[26]  Douglas M. Hawkins,et al.  A variance-stabilizing transformation for gene-expression microarray data , 2002, ISMB.

[27]  Robert Tibshirani,et al.  An Introduction to the Bootstrap , 1994 .

[28]  P. Broberg Statistical methods for ranking differentially expressed genes , 2003, Genome Biology.

[29]  X. Cui,et al.  Improved statistical tests for differential gene expression by shrinking variance components estimates. , 2005, Biostatistics.

[30]  D. Bickel Microarray Gene Expression Analysis: Data Transformation and Multiple-Comparison Bootstrapping , 2002 .

[31]  Ingrid Lönnstedt Replicated microarray data , 2001 .

[32]  John D. Storey,et al.  Statistical significance for genomewide studies , 2003, Proceedings of the National Academy of Sciences of the United States of America.

[33]  John D. Storey,et al.  Empirical Bayes Analysis of a Microarray Experiment , 2001 .

[34]  J. S. Rao,et al.  Spike and Slab Gene Selection for Multigroup Microarray Data , 2005 .

[35]  Gordon K Smyth,et al.  Linear Models and Empirical Bayes Methods for Assessing Differential Expression in Microarray Experiments , 2004, Statistical applications in genetics and molecular biology.

[36]  Hua Xu,et al.  R package MVR for Joint Adaptive Mean-Variance Regularization and Variance Stabilization. , 2011, Proceedings. American Statistical Association. Annual Meeting.

[37]  Jianqing Fan,et al.  Sure independence screening for ultrahigh dimensional feature space , 2006, math/0612857.

[38]  David M. Rocke,et al.  A Model for Measurement Error for Gene Expression Arrays , 2001, J. Comput. Biol..

[39]  J. Neyman,et al.  INADMISSIBILITY OF THE USUAL ESTIMATOR FOR THE MEAN OF A MULTIVARIATE NORMAL DISTRIBUTION , 2005 .

[40]  Jae K. Lee,et al.  Local-pooled-error test for identifying differentially expressed genes with a small number of replicated microarrays , 2003, Bioinform..

[41]  Tiejun Tong,et al.  Optimal Shrinkage Estimation of Variances With Applications to Microarray Data Analysis , 2007 .

[42]  Yongchao Ge Resampling-based Multiple Testing for Microarray Data Analysis , 2003 .

[43]  Robert Tibshirani,et al.  Estimating the number of clusters in a data set via the gap statistic , 2000 .

[44]  John D. Storey The optimal discovery procedure: a new approach to simultaneous significance testing , 2007 .

[45]  Terence P. Speed,et al.  A comparison of normalization methods for high density oligonucleotide array data based on variance and bias , 2003, Bioinform..

[46]  S. Dudoit,et al.  STATISTICAL METHODS FOR IDENTIFYING DIFFERENTIALLY EXPRESSED GENES IN REPLICATED cDNA MICROARRAY EXPERIMENTS , 2002 .

[47]  Pierre Baldi,et al.  A Bayesian framework for the analysis of microarray expression data: regularized t -test and statistical inferences of gene changes , 2001, Bioinform..

[48]  Phillip I. Good,et al.  Extensions Of The Concept Of Exchangeability And Their Applications , 2002 .