论文信息 - Use of Line Spectral Frequencies for Emotion Recognition from Speech

Use of Line Spectral Frequencies for Emotion Recognition from Speech

We propose the use of the line spectral frequency (LSF) features for emotion recognition from speech, which have not been been previously employed for emotion recognition to the best of our knowledge. Spectral features such as mel-scaled cepstral coefficients have already been successfully used for the parameterization of speech signals for emotion recognition. The LSF features also offer a spectral representation for speech, moreover they carry intrinsic information on the formant structure as well, which are related to the emotional state of the speaker [4]. We use the Gaussian mixture model (GMM) classifier architecture, that captures the static color of the spectral features. Experimental studies performed over the Berlin Emotional Speech Database and the FAU Aibo Emotion Corpus demonstrate that decision fusion configurations with LSF features bring a consistent improvement over the MFCC based emotion classification rates.

A. Tanju Erdem | Engin Erzin | Çigdem Eroglu Erdem | Elif Bozkurt

[1] Frank A. Andrews,et al. Measurements of radiation impedance , 1975 .

[2] R.W. Morris,et al. Modification of formants in the line spectrum domain , 2002, IEEE Signal Processing Letters.

[3] F. Itakura. Line spectrum representation of linear predictor coefficients of speech signals , 1975 .

[4] Björn W. Schuller,et al. Frame vs. Turn-Level: Emotion Recognition from Speech Considering Static and Dynamic Processing , 2007, ACII.

[5] Wei Zhang,et al. EM algorithms of Gaussian mixture model and hidden Markov model , 2001, Proceedings 2001 International Conference on Image Processing (Cat. No.01CH37205).

[6] A. Murat Tekalp,et al. Multimodal speaker identification using an adaptive classifier cascade based on modality reliability , 2005, IEEE Transactions on Multimedia.

[7] Klaus R. Scherer,et al. Emotion dimensions and formant position , 2009, INTERSPEECH.

[8] Björn Schuller,et al. Emotion Recognition in the Noise Applying Large Acoustic Feature Sets , 2006 .

[9] Stefan Steidl,et al. Automatic classification of emotion related user states in spontaneous children's speech , 2009 .

[10] Björn W. Schuller,et al. The INTERSPEECH 2009 emotion challenge , 2009, INTERSPEECH.

[11] Astrid Paeschke,et al. A database of German emotional speech , 2005, INTERSPEECH.

[12] A. Tanju Erdem,et al. Improving automatic emotion recognition from speech signals , 2009, INTERSPEECH.