Identification of MicroRNA Precursors with Support Vector Machine and String Kernel

MicroRNAs (miRNAs) are one family of short (21–23 nt) regulatory non-coding RNAs processed from long (70–110 nt) miRNA precursors (pre-miRNAs). Identifying true and false precursors plays an important role in computational identification of miRNAs. Some numerical features have been extracted from precursor sequences and their secondary structures to suit some classification methods; however, they may lose some usefully discriminative information hidden in sequences and structures. In this study, pre-miRNA sequences and their secondary structures are directly used to construct an exponential kernel based on weighted Levenshtein distance between two sequences. This string kernel is then combined with support vector machine (SVM) for detecting true and false pre-miRNAs. Based on 331 training samples of true and false human pre-miRNAs, 2 key parameters in SVM are selected by 5-fold cross validation and grid search, and 5 realizations with different 5-fold partitions are executed. Among 16 independent test sets from 3 human, 8 animal, 2 plant, 1 virus, and 2 artificially false human pre-miRNAs, our method statistically outperforms the previous SVM-based technique on 11 sets, including 3 human, 7 animal, and 1 false human pre-miRNAs. In particular, pre-miRNAs with multiple loops that were usually excluded in the previous work are correctly identified in this study with an accuracy of 92.66%.

[1]  C. Burge,et al.  The microRNAs of Caenorhabditis elegans. , 2003, Genes & development.

[2]  Sam Griffiths-Jones,et al.  The microRNA Registry , 2004, Nucleic Acids Res..

[3]  Mihaela Zavolan,et al.  Identification of Clustered Micrornas Using an Ab Initio Prediction Method , 2022 .

[4]  D. Bartel MicroRNAs Genomics, Biogenesis, Mechanism, and Function , 2004, Cell.

[5]  Vladimir N. Vapnik,et al.  The Nature of Statistical Learning Theory , 2000, Statistics for Engineering and Information Science.

[6]  Michael J. Fischer,et al.  The String-to-String Correction Problem , 1974, JACM.

[7]  Vladimir Vapnik,et al.  Statistical learning theory , 1998 .

[8]  David G. Stork,et al.  Pattern Classification , 1973 .

[9]  Todd A. Anderson,et al.  Computational identification of microRNAs and their targets , 2006, Comput. Biol. Chem..

[10]  Torbjørn Rognes,et al.  Computational Prediction of MicroRNAs Encoded in Viral and Other Genomes , 2006, Journal of biomedicine & biotechnology.

[11]  Baohong Zhang,et al.  MicroRNAs and their regulatory roles in animals and plants , 2007, Journal of cellular physiology.

[12]  Leo Breiman,et al.  Random Forests , 2001, Machine Learning.

[13]  Fei Li,et al.  MicroRNA identification based on sequence and structure alignment , 2005, Bioinform..

[14]  Jason Weston,et al.  Mismatch string kernels for discriminative protein classification , 2004, Bioinform..

[15]  Reiji Teramoto,et al.  Prediction of siRNA functionality using generalized string kernel and support vector machine , 2005, FEBS letters.

[16]  Stijn van Dongen,et al.  miRBase: microRNA sequences, targets and gene nomenclature , 2005, Nucleic Acids Res..

[17]  C. Burge,et al.  Vertebrate MicroRNA Genes , 2003, Science.

[18]  Jianhua Xu,et al.  Kernels based on weighted Levenshtein distance , 2004, 2004 IEEE International Joint Conference on Neural Networks (IEEE Cat. No.04CH37541).

[19]  King-Sun Fu,et al.  Syntactic Pattern Recognition And Applications , 1968 .

[20]  James R. Brown,et al.  A computational view of microRNAs and their targets. , 2005, Drug discovery today.

[21]  V. Kim,et al.  MicroRNA maturation: stepwise processing and subcellular localization , 2002, The EMBO journal.

[22]  P. Rouzé,et al.  Detection of 91 potential conserved plant microRNAs in Arabidopsis thaliana and Oryza sativa identifies important target genes. , 2004, Proceedings of the National Academy of Sciences of the United States of America.

[23]  Fei Li,et al.  Classification of real and pseudo microRNA precursors using local structure-sequence features and support vector machine , 2005, BMC Bioinformatics.

[24]  Peng Jiang,et al.  MiPred: classification of real and pseudo microRNA precursors using random forest prediction model with combined features , 2007, Nucleic Acids Res..

[25]  D. Bartel,et al.  Computational identification of plant microRNAs and their targets, including a stress-induced miRNA. , 2004, Molecular cell.

[26]  Fang Chen,et al.  Gene expression regulators —MicroRNAs , 2005 .

[27]  Vladimir N. Vapnik,et al.  The Nature of Statistical Learning Theory, Second Edition , 2000, Statistics for Engineering and Information Science.

[28]  Yuichiro Watanabe,et al.  Arabidopsis micro-RNA biogenesis through Dicer-like 1 protein functions. , 2004, Proceedings of the National Academy of Sciences of the United States of America.

[29]  G. Rubin,et al.  Computational identification of Drosophila microRNA genes , 2003, Genome Biology.

[30]  Walter Fontana,et al.  Fast folding and comparison of RNA secondary structures , 1994 .

[31]  Vladimir I. Levenshtein,et al.  Binary codes capable of correcting deletions, insertions, and reversals , 1965 .