论文信息 - IPA: improved phone modelling with recurrent neural networks

IPA: improved phone modelling with recurrent neural networks

This paper describes phone modelling improvements to the hybrid connectionist-hidden Markov model speech recognition system developed at Cambridge University. These improvements are applied to phone recognition from the TIMIT task and word recognition from the Wall Street Journal (WSJ) task. A recurrent net is used to map acoustic vectors to posterior probabilities of phone classes. The maximum likelihood phone or word string is then extracted using Markov models. The paper describes three improvements: connectionist model merging; explicit presentation of acoustic context; and improved duration modelling. The first is shown to provide a significant improvement in the TIMIT phone recognition rate and all three provide an improvement in the WSJ word recognition rate.<<ETX>>

[1] Anthony J. Robinson,et al. An application of recurrent nets to phone probability estimation , 1994, IEEE Trans. Neural Networks.

[2] Yochai Konig,et al. A neural network based, speaker independent, large vocabulary, continuous speech recognition system: the WERNICKE project , 1993, EUROSPEECH.

[3] Wray L. Buntine,et al. Learning classification trees , 1992 .

[4] Janet M. Baker,et al. The Design for the Wall Street Journal-based CSR Corpus , 1992, HLT.

[5] David H. Wolpert,et al. Stacked generalization , 1992, Neural Networks.

[6] Jean-Luc Gauvain,et al. High performance speaker-independent phone recognition using CDHMM , 1993, EUROSPEECH.

[7] T. H. Crystal,et al. Segmental durations in connected speech signals , 1981 .

[8] H Hermansky,et al. Perceptual linear predictive (PLP) analysis of speech. , 1990, The Journal of the Acoustical Society of America.