Reference-free cell mixture adjustments in analysis of DNA methylation data

Motivation: Recently there has been increasing interest in the effects of cell mixture on the measurement of DNA methylation, specifically the extent to which small perturbations in cell mixture proportions can register as changes in DNA methylation. A recently published set of statistical methods exploits this association to infer changes in cell mixture proportions, and these methods are presently being applied to adjust for cell mixture effect in the context of epigenome-wide association studies. However, these adjustments require the existence of reference datasets, which may be laborious or expensive to collect. For some tissues such as placenta, saliva, adipose or tumor tissue, the relevant underlying cell types may not be known. Results: We propose a method for conducting epigenome-wide association studies analysis when a reference dataset is unavailable, including a bootstrap method for estimating standard errors. We demonstrate via simulation study and several real data analyses that our proposed method can perform as well as or better than methods that make explicit use of reference datasets. In particular, it may adjust for detailed cell type differences that may be unavailable even in existing reference datasets. Availability and implementation: Software is available in the R package RefFreeEWAS. Data for three of four examples were obtained from Gene Expression Omnibus (GEO), accession numbers GSE37008, GSE42861 and GSE30601, while reference data were obtained from GEO accession number GSE39981. Contact: andres.houseman@oregonstate.edu Supplementary information: Supplementary data are available at Bioinformatics online.

[1]  Andrew E. Teschendorff,et al.  Independent surrogate variable analysis to deconvolve confounding factors in large-scale microarray profiling studies , 2011, Bioinform..

[2]  J. Rinn,et al.  DNA methylation and epigenetic control of cellular differentiation , 2010, Cell cycle.

[3]  B. Teh,et al.  Methylation Subtypes and Large-Scale Epigenetic Alterations in Gastric Cancer , 2012, Science Translational Medicine.

[4]  Eldon Emberly,et al.  Factors underlying variable DNA methylation in a human community cohort , 2012, Proceedings of the National Academy of Sciences.

[5]  Kristian Helin,et al.  Genome-wide mapping of Polycomb target genes unravels their roles in cell fate transitions. , 2006, Genes & development.

[6]  Zohar Yakhini,et al.  Polycomb-mediated methylation on Lys27 of histone H3 pre-marks genes for de novo methylation in cancer , 2007, Nature Genetics.

[7]  Margaret R Karagas,et al.  Blood-based profiles of DNA methylation predict the underlying distribution of cell types , 2013, Epigenetics.

[8]  John D. Storey,et al.  Strong control, conservative point estimation and simultaneous conservative consistency of false discovery rates: a unified approach , 2004 .

[9]  C. Bock Analysing and interpreting DNA methylation data , 2012, Nature Reviews Genetics.

[10]  Devin C. Koestler,et al.  DNA methylation arrays as surrogate measures of cell mixture distribution , 2012, BMC Bioinformatics.

[11]  Margaret R Karagas,et al.  Peripheral Blood Immune Cell Methylation Profiles Are Associated with Nonhematopoietic Cancers , 2012, Cancer Epidemiology, Biomarkers & Prevention.

[12]  Henriette O'Geen,et al.  Suz12 binds to silenced regions of the genome in a cell-type-specific manner. , 2006, Genome research.

[13]  Vilmundur Gudnason,et al.  Heterogeneity in White Blood Cells Has Potential to Confound DNA Methylation Measurements , 2012, PloS one.

[14]  G. Natoli Maintaining cell identity through global control of genomic organization. , 2010, Immunity.

[15]  J. Kere,et al.  Differential DNA Methylation in Purified Human Blood Cells: Implications for Cell Lineage and Studies on Disease Susceptibility , 2012, PloS one.

[16]  Martin J. Aryee,et al.  Epigenome-wide association data implicate DNA methylation as an intermediary of genetic risk in Rheumatoid Arthritis , 2013, Nature Biotechnology.

[17]  Sven Olek,et al.  DNA Methylation Analysis as a Tool for Cell Typing , 2006, Epigenetics.

[18]  Megan F. Cole,et al.  Control of Developmental Regulators by Polycomb in Human Embryonic Stem Cells , 2006, Cell.

[19]  Irving L. Weissman,et al.  A comprehensive methylome map of lineage commitment from hematopoietic progenitors , 2010, Nature.

[20]  Cheng Li,et al.  Adjusting batch effects in microarray expression data using empirical Bayes methods. , 2007, Biostatistics.

[21]  Devin C Koestler,et al.  Infant growth restriction is associated with distinct patterns of DNA methylation in human placentas , 2011, Epigenetics.

[22]  Gordon K Smyth,et al.  Linear Models and Empirical Bayes Methods for Assessing Differential Expression in Microarray Experiments , 2004, Statistical applications in genetics and molecular biology.

[23]  A. Buja,et al.  Remarks on Parallel Analysis. , 1992, Multivariate behavioral research.

[24]  John D. Storey,et al.  Capturing Heterogeneity in Gene Expression Studies by Surrogate Variable Analysis , 2007, PLoS genetics.

[25]  Rondi A. Butler,et al.  Peripheral blood DNA methylation profiles are indicative of head and neck squamous cell carcinoma: An epigenome-wide association study , 2012, Epigenetics.

[26]  Huidong Shi,et al.  A Genome-Wide Methylation Study on Essential Hypertension in Young African American Males , 2013, PloS one.