On multiple‐testing correction in genome‐wide association studies

The interpretation of the results of large association studies encompassing much or all of the human genome faces the fundamental statistical problem that a correspondingly large number of single nucleotide polymorphisms markers will be spuriously flagged as significant. A common method of dealing with these false positives is to raise the significance level for the individual tests for association of each marker. Any such adjustment for multiple testing is ultimately based on a more or less precise estimate for the actual overall type I error probability. We estimate this probability for association tests for correlated markers and show that it depends in a nonlinear way on the significance level for the individual tests. This dependence of the effective number of tests is not taken into account by existing multiple‐testing corrections, leading to widely overestimated results. We demonstrate a simple correction for multiple testing, which can easily be calculated from the pairwise correlation and gives far more realistic estimates for the effective number of tests than previous formulae. The calculation is considerably faster than with other methods and hence applicable on a genome‐wide scale. The efficacy of our method is shown on a constructed example with highly correlated markers as well as on real data sets, including a full genome scan where a conservative estimate only 8% above the permutation estimate is obtained in about 1% of computation time. As the calculation is based on pairwise correlations between markers, it can be performed at the stage of study design using public databases. Genet. Epidemiol. 2008. 2008 Wiley‐Liss, Inc.

[1]  H. Cramér Mathematical methods of statistics , 1947 .

[2]  Z. Šidák Rectangular Confidence Regions for the Means of Multivariate Normal Distributions , 1967 .

[3]  J. Cheverud,et al.  A simple correction for multiple comparisons in interval mapping genome scans , 2001, Heredity.

[4]  John D. Storey,et al.  Statistical significance for genomewide studies , 2003, Proceedings of the National Academy of Sciences of the United States of America.

[5]  D. Zaykin,et al.  Effect of Two- and Three-Locus Linkage Disequilibrium on the Power to Detect Marker/Phenotype Associations , 2004, Genetics.

[6]  D. Nyholt A simple correction for multiple testing for single-nucleotide polymorphisms in linkage disequilibrium with each other. , 2004, American journal of human genetics.

[7]  Frank Dudbridge,et al.  Efficient computation of significance levels for multiple associations in large studies of correlated data, including genomewide association studies. , 2004, American journal of human genetics.

[8]  Dale R Nyholt,et al.  Evaluation of Nyholt’s Procedure for Multiple Testing Correction – Author’s Reply , 2005, Human Heredity.

[9]  Frank Dudbridge,et al.  Evaluation of Nyholt’s Procedure for Multiple Testing Correction , 2005, Human Heredity.

[10]  J. Terwilliger,et al.  An utter refutation of the ‘Fundamental Theorem of the HapMap’ , 2006, European Journal of Human Genetics.

[11]  M. Boehnke,et al.  So many correlated tests, so little time! Rapid adjustment of P values for multiple correlated tests. , 2007, American journal of human genetics.

[12]  S. E. Ahmed,et al.  Handbook of Statistical Distributions with Applications , 2007, Technometrics.

[13]  Genomewide association study of 14 , 000 cases of seven common diseases and 3 , 000 shared controls Supplementary Information , 2022 .