Local feature based gender independent bangla ASR

This paper presents automatic speech recognition (ASR) for Bangla (widely used as Bengali) by suppressing the speaker gender types based on local features extracted from an input speech. Speaker-specific characteristics play an important role on the performance of Bangla automatic speech recognition (ASR). Gender factor shows adverse effect in the classifier while recognizing a speech by an opposite gender, such as, training a classifier by male but testing is done by female or vice-versa. To obtain a robust ASR system in practice it is necessary to invent a system that incorporates gender independent effect for particular gender. In this paper, we have proposed a Gender-Independent technique for ASR that focused on a gender factor. The proposed method trains the classifier with the both types of gender, male and female, and evaluates the classifier for the male and female. For the experiments, we have designed a medium size Bangla (widely known as Bengali) speech corpus for both the male and female. The proposed system has showed a significant improvement of word correct rates, word accuracies and sentence correct rates in comparison with the method that suffers from gender effects using. Moreover, it provides the highest level recognition performance by taking a fewer mixture component in hidden Markov model (HMMs).

[1]  Mumit Khan,et al.  Isolated and continuous bangla speech recognition: implementation, performance and application perspective , 2007 .

[2]  Ghulam Muhammad,et al.  Automatic speech recognition for Bangla digits , 2009, 2009 12th International Conference on Computers and Information Technology.

[3]  Syed Akhter Hossain,et al.  Bangla Vowel Characterization Based on Analysis by Synthesis , 2007 .

[4]  Mohammed Rokibul Alam Kotwal,et al.  Gender independent Bangla automatic speech recognition , 2012, 2012 International Conference on Informatics, Electronics & Vision (ICIEV).

[5]  Colin P. Masica The Indo-Aryan Languages , 1991 .

[6]  Satoshi Nakamura,et al.  ATR Parallel Decoding Based Speech Recognition System Robust to Noise and Speaking Styles , 2006, IEICE Trans. Inf. Syst..

[7]  A. Black,et al.  1 Experiments with Unit Selection Speech Databases for Indian Languages , 2003 .

[8]  Tsuneo Nitta Feature extraction for speech recognition based on orthogonal acoustic-feature planes and LDA , 1999, 1999 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings. ICASSP99 (Cat. No.99CH36258).