论文信息 - Speech Recognition Using Energy Parameters to Classify Syllables in the Spanish Language

Speech Recognition Using Energy Parameters to Classify Syllables in the Spanish Language

This paper presents an approach for the automatic speech re-cognition using syllabic units. Its segmentation is based on using the Short-Term Total Energy Function (STTEF) and the Energy Function of the High Frequency (ERO parameter) higher than 3,5 KHz of the speech signal. Training for the classification of the syllables is based on ten related Spanish language rules for syllable splitting. Recognition is based on a Continuous Density Hidden Markov Models and the bigram model language. The approach was tested using two voice corpus of natural speech, one constructed for researching in our laboratory (experimental) and the other one, the corpus Latino40 commonly used in speech researches. The use of ERO parameter increases speech recognition by 5% when compared with recognition using STTEF in discontinuous speech and improved more than 1.5% in continuous speech with three states. When the number of states is incremented to five, the recognition rate is improved proportionally to 97.5% for the discontinuous speech and to 80.5% for the continuous one.

Edgardo Manuel Felipe Riverón | José Luis Oropeza Rodríguez | Sergio Suárez Guerra | Jesús Nazuno

[1] Sergio Suárez Guerra,et al. Pruebas y utilización de un sistema de reconocimiento del habla basado en sílabas con un vocabulario pequeño , 2003 .

[2] Jeff A. Bilmes,et al. A gentle tutorial of the em algorithm and its application to parameter estimation for Gaussian mixture and hidden Markov models , 1998 .

[3] Biing-Hwang Juang,et al. Fundamentals of speech recognition , 1993, Prentice Hall signal processing series.

[4] M. Meyerhoff,et al. Working papers in linguistics , 1994 .

[5] Steven Greenberg,et al. Incorporating information from syllable-length time scales into automatic speech recognition , 1998, Proceedings of the 1998 IEEE International Conference on Acoustics, Speech and Signal Processing, ICASSP '98 (Cat. No.98CH36181).

[6] Hidehito Hoshi. UCI working papers in linguistics , 1996 .

[7] Jesus Savage-Carmona,et al. A hybrid system with symbolic AI and statistical methods for speech recognition , 1996 .

[8] João Paulo da Silva Neto,et al. Combination of acoustic models in continuous speech recognition hybrid systems , 2000, INTERSPEECH.

[9] Steven Greenberg,et al. Integrating syllable boundary information into speech recognition , 1997, 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing.

[10] João Paulo da Silva Neto,et al. Syllable onset detection applied to the portuguese language , 1999, EUROSPEECH.