论文信息 - A real-time Japanese broadcast news closed-captioning system

A real-time Japanese broadcast news closed-captioning system

This paper describes a collaboration between Bell Labs and NHK (Japan Broadcasting Corp.) STRL to develop a real-time large vocabulary speech recognition system for live closed-captioning of NHK news programs. Bell Labs broadcast news recognition engine consists of a two-pass decoder using bigram language models (LM) and right biphone models during the first pass, and trigram LM with within-word triphone models in the second pass. Various pruning strategies are used to achieve real time decoding, together with a noise compensation procedure aimed at improving recognition on noisy segments of the program. The system operates in a real-time mode and delivers less than 2% of word error rate (WER) on studio news conditions and about 5% of WER on noisy news and reporter speech when evaluated on a real broadcast news program.

[1] Richard M. Schwartz,et al. Single-tree method for grammar-directed search , 1999, 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No.99CH36258).

[2] Kazuo Onoe,et al. Time dependent language model for broadcast news transcription and its post-correction , 1998, ICSLP.

[3] Olivier Siohan,et al. Sequential noise estimation with optimal forgetting for robust speech recognition , 2001, 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.01CH37221).

[4] Wu Chou,et al. Robust decision tree state tying for continuous speech recognition , 2000, IEEE Trans. Speech Audio Process..

[5] R. A. Sukkar,et al. Variable threshold vector quantization for reduced continuous density likelihood computation in speech recognition , 1997, 1997 IEEE Workshop on Automatic Speech Recognition and Understanding Proceedings.

[6] Olivier Siohan,et al. A NEW VERIFICATION-BASED FAST MATCH APPROACH TO LARGE VOCABULARY CONSTINUOUS SPEECH RECOGNITION , 2001 .

[7] Shoei Sato,et al. Progressive 2-pass decoder for real-time broadcast news captioning , 2000, 2000 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No.00CH37100).

[8] Qiru Zhou,et al. An approach to continuous speech recognition based on layered self-adjusting decoding graph , 1997, 1997 IEEE International Conference on Acoustics, Speech, and Signal Processing.