论文信息 - The Protein Naming Utility: a rules database for protein nomenclature

The Protein Naming Utility: a rules database for protein nomenclature

Generation of syntactically correct and unambiguous names for proteins is a challenging, yet vital task for functional annotation processes. Proteins are often named based on homology to known proteins, many of which have problematic names. To address the need to generate high-quality protein names, and capture our significant experience correcting protein names manually, we have developed the Protein Naming Utility (PNU, http://www.jcvi.org/pn-utility). The PNU is a web-based database for storing and applying naming rules to identify and correct syntactically incorrect protein names, or to replace synonyms with their preferred name. The PNU allows users to generate and manage collections of naming rules, optionally building upon the growing body of rules generated at the J. Craig Venter Institute (JCVI). Since communities often enforce disparate conventions for naming proteins, the PNU supports grouping rules into user-managed collections. Users can check their protein names against a selected PNU rule collection, generating both statistics and corrected names. The PNU can also be used to correct GenBank table files prior to submission to GenBank. Currently, the database features 3080 manual rules that have been entered by JCVI Bioinformatics Analysts as well as 7458 automatically imported names.

[1] Kara Dolinski,et al. Saccharomyces genome database: Underlying principles and organisation , 2004, Briefings Bioinform..

[2] Hongfang Liu,et al. BioThesaurus: a web-based thesaurus of protein and gene names , 2006, Bioinform..

[3] Sue Povey,et al. The HUGO Gene Nomenclature Database, 2006 updates , 2005, Nucleic Acids Res..

[4] Judith A. Blake,et al. The Mouse Genome Database (MGD): mouse biology and model systems , 2007, Nucleic Acids Res..

[5] Andrew G. McDonald,et al. ExplorEnz: the primary source of the IUBMB enzyme list , 2008, Nucleic Acids Res..

[6] Melinda R. Dwinell,et al. The Rat Genome Database 2009: variation, ontologies and pathways , 2008, Nucleic Acids Res..

[7] Ralf Zimmer,et al. Gene and protein nomenclature in public databases , 2006, BMC Bioinformatics.

[8] Madeline A. Crosby,et al. FlyBase: genes and gene models , 2004, Nucleic Acids Res..

[9] The UniProt Consortium,et al. The Universal Protein Resource (UniProt) 2009 , 2008, Nucleic Acids Res..