Speech Emotion Recognition under White Noise

被引:27
|
作者
Huang, Chengwei [1 ]
Chen, Guoming [1 ]
Yu, Hua [1 ]
Bao, Yongqiang [2 ]
Zhao, Li [1 ]
机构
[1] Southeast Univ, Sch Informat Sci & Engn, Nanjing 210096, Jiangsu, Peoples R China
[2] Nanjing Inst Technol, Sch Commun Engn, Nanjing 211167, Jiangsu, Peoples R China
关键词
speech emotion recognition; speech enhancement; emotion model; Gaussian mixture model; ENHANCEMENT; AUDIO;
D O I
10.2478/aoa-2013-0054
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
Speaker's emotional states are recognized from speech signal with Additive white Gaussian noise (AWGN). The influence of white noise on a typical emotion recogniztion system is studied. The emotion classifier is implemented with Gaussian mixture model (GMM). A Chinese speech emotion database is used for training and testing, which includes nine emotion classes (e.g. happiness, sadness, anger, surprise, fear, anxiety, hesitation, confidence and neutral state). Two speech enhancement algorithms are introduced for improved emotion classification. In the experiments, the Gaussian mixture model is trained on the clean speech data, while tested under AWGN with various signal to noise ratios (SNRs). The emotion class model and the dimension space model are both adopted for the evaluation of the emotion recognition system. Regarding the emotion class model, the nine emotion classes are classified. Considering the dimension space model, the arousal dimension and the valence dimension are classified into positive regions or negative regions: The experimental results show that the speech enhancement algorithms constantly improve the performance of our emotion recognition system under various SNRs, and the positive emotions are more likely to be miss-classified as negative emotions under white noise environment.
引用
收藏
页码:457 / 463
页数:7
相关论文
共 50 条
  • [1] Emotion Recognition from Speech under Environmental Noise Conditions using Wavelet Decomposition
    Vasquez-Correa, J. C.
    Garcia, N.
    Orozco-Arroyave, J. R.
    Arias-Londono, J. D.
    Vargas-Bonilla, J. F.
    Noeth, Elmar
    [J]. 49TH ANNUAL IEEE INTERNATIONAL CARNAHAN CONFERENCE ON SECURITY TECHNOLOGY (ICCST), 2015, : 247 - 252
  • [2] Improving Noise Robustness of Speech Emotion Recognition System
    Juszkiewicz, Lukasz
    [J]. INTELLIGENT DISTRIBUTED COMPUTING VII, 2014, 511 : 223 - 232
  • [3] Speech Emotion Recognition
    Lalitha, S.
    Madhavan, Abhishek
    Bhushan, Bharath
    Saketh, Srinivas
    [J]. 2014 INTERNATIONAL CONFERENCE ON ADVANCES IN ELECTRONICS, COMPUTERS AND COMMUNICATIONS (ICAECC), 2014,
  • [4] MANDARIN AUDIO-VISUAL SPEECH RECOGNITION WITH EFFECTS TO THE NOISE AND EMOTION
    Pao, Tsang-Long
    Liao, Wen-Yuan
    Chen, Yu-Te
    Wu, Tsan-Nung
    [J]. INTERNATIONAL JOURNAL OF INNOVATIVE COMPUTING INFORMATION AND CONTROL, 2010, 6 (02): : 711 - 723
  • [5] Spectral and Cepstral Audio Noise Reduction Techniques in Speech Emotion Recognition
    Pohjalainen, Jouni
    Ringeval, Fabien
    Zhang, Zixing
    Schuller, Bjoern
    [J]. MM'16: PROCEEDINGS OF THE 2016 ACM MULTIMEDIA CONFERENCE, 2016, : 670 - 674
  • [6] Robust Speech Emotion Recognition under Different Encoding Conditions
    Oates, Christopher
    Triantafyllopoulos, Andreas
    Steiner, Ingmar
    Schuller, Bjoern
    [J]. INTERSPEECH 2019, 2019, : 3935 - 3939
  • [7] Feature fusion methods research based on deep belief networks for speech emotion recognition under noise condition
    Yongming Huang
    Kexin Tian
    Ao Wu
    Guobao Zhang
    [J]. Journal of Ambient Intelligence and Humanized Computing, 2019, 10 : 1787 - 1798
  • [8] Feature fusion methods research based on deep belief networks for speech emotion recognition under noise condition
    Huang, Yongming
    Tian, Kexin
    Wu, Ao
    Zhang, Guobao
    [J]. JOURNAL OF AMBIENT INTELLIGENCE AND HUMANIZED COMPUTING, 2019, 10 (05) : 1787 - 1798
  • [9] Research on Robustness of Emotion Recognition Under Environmental Noise Conditions
    Huang, Yongming
    Xiao, Jing
    Tian, Kexin
    Wu, Ao
    Zhang, Guobao
    [J]. IEEE ACCESS, 2019, 7 : 142009 - 142021
  • [10] Towards Noise Robust Speech Emotion Recognition Using Dynamic Layer Customization
    Wilf, Alex
    Provost, Emily Mower
    [J]. 2021 9TH INTERNATIONAL CONFERENCE ON AFFECTIVE COMPUTING AND INTELLIGENT INTERACTION (ACII), 2021,