Towards de novo identification of metabolites by analyzing tandem mass spectra

MOTIVATION Mass spectrometry is among the most widely used technologies in proteomics and metabolomics. Being a high-throughput method, it produces large amounts of data that necessitates an automated analysis of the spectra. Clearly, database search methods for protein analysis can easily be adopted to analyze metabolite mass spectra. But for metabolites, de novo interpretation of spectra is even more important than for protein data, because metabolite spectra databases cover only a small fraction of naturally occurring metabolites: even the model plant Arabidopsis thaliana has a large number of enzymes whose substrates and products remain unknown. The field of bio-prospection searches biologically diverse areas for metabolites which might serve as pharmaceuticals. De novo identification of metabolite mass spectra requires new concepts and methods since, unlike proteins, metabolites possess a non-linear molecular structure. RESULTS In this work, we introduce a method for fully automated de novo identification of metabolites from tandem mass spectra. Mass spectrometry data is usually assumed to be insufficient for identification of molecular structures, so we want to estimate the molecular formula of the unknown metabolite, a crucial step for its identification. The method first calculates all molecular formulas that explain the parent peak mass. Then, a graph is build where vertices correspond to molecular formulas of all peaks in the fragmentation mass spectra, whereas edges correspond to hypothetical fragmentation steps. Our algorithm afterwards calculates the maximum scoring subtree of this graph: each peak in the spectra must be scored at most once, so the subtree shall contain only one explanation per peak. Unfortunately, finding this subtree is NP-hard. We suggest three exact algorithms (including one fixed parameter tractable algorithm) as well as two heuristics to solve the problem. Tests on real mass spectra show that the FPT algorithm and the heuristics solve the problem suitably fast and provide excellent results: for all 32 test compounds the correct solution was among the top five suggestions, for 26 compounds the first suggestion of the exact algorithm was correct. AVAILABILITY http://www.bio.inf.uni-jena.de/tandemms

[1]  Oliver Fiehn,et al.  Metabolomic database annotations via query of elemental compositions: Mass accuracy is insufficient even at less than 1 ppm , 2006, BMC Bioinformatics.

[2]  J. Kruskal On the shortest spanning subtree of a graph and the traveling salesman problem , 1956 .

[3]  Oliver Fiehn,et al.  Seven Golden Rules for heuristic filtering of molecular formulas obtained by accurate mass spectrometry , 2007, BMC Bioinformatics.

[4]  S. A. McLuckey,et al.  Collision-induced dissociation (CID) of peptides and proteins. , 2005, Methods in enzymology.

[5]  Zsuzsanna Lipták,et al.  A Fast and Simple Algorithm for the Money Changing Problem , 2007, Algorithmica.

[6]  Antony Williams,et al.  Applications of computer software for the interpretation and management of mass spectrometry data in pharmaceutical science. , 2002, Current topics in medicinal chemistry.

[7]  Ming-Yang Kao,et al.  A dynamic programming approach to de novo peptide sequencing via tandem mass spectrometry , 2000, SODA '00.

[8]  A. Masselot,et al.  Assessing peptide de novo sequencing algorithms performance on large and diverse data sets , 2007, Proteomics.

[9]  Ludger Wessjohann,et al.  Profiling of Arabidopsis Secondary Metabolites by Capillary Liquid Chromatography Coupled to Electrospray Ionization Quadrupole Time-of-Flight Mass Spectrometry1 , 2004, Plant Physiology.

[10]  Dimitrios M. Thilikos,et al.  Invitation to fixed-parameter algorithms , 2007, Comput. Sci. Rev..

[11]  Juho Rousu,et al.  Ab Initio Prediction of Molecular Fragments from Tandem Mass Spectrometry Data , 2006, German Conference on Bioinformatics.

[12]  Michael R. Fellows,et al.  Sharp Tractability Borderlines for Finding Connected Motifs in Vertex-Colored Graphs , 2007, ICALP.

[13]  Fred W. McLafferty,et al.  BenchTop/PBM : Wiley registry of mass spectral data 7th edition, with NIST 2002 , 2003 .

[14]  The Arabidopsis Genome Initiative Analysis of the genome sequence of the flowering plant Arabidopsis thaliana , 2000, Nature.

[15]  J. Gershenzon,et al.  The secondary metabolism of Arabidopsis thaliana: growing like a weed. , 2005, Current opinion in plant biology.

[16]  J. K. Senior Partitions and Their Representative Graphs , 1951 .

[17]  Zsuzsanna Lipták,et al.  Decomposing Metabolomic Isotope Patterns , 2006, WABI.

[18]  Roded Sharan,et al.  Efficient Algorithms for Detecting Signaling Pathways in Protein Interaction Networks , 2006, J. Comput. Biol..

[19]  David S. Johnson,et al.  Computers and Intractability: A Guide to the Theory of NP-Completeness , 1978 .

[20]  Kiyoko F. Aoki-Kinoshita,et al.  From genomics to chemical genomics: new developments in KEGG , 2005, Nucleic Acids Res..