A Cross-Corpus Speech-Based Analysis of Escalating Negative Interactions

被引:3
|
作者
Lefter, Iulia [1 ]
Baird, Alice [2 ]
Stappen, Lukas [2 ]
Schuller, Bjorn W. [2 ,3 ]
机构
[1] Delft Univ Technol, Dept Multiactor Syst, Delft, Netherlands
[2] Univ Augsburg, Chair Embedded Intelligence Hlth Care & Wellbeing, Augsburg, Germany
[3] Imperial Coll London, Grp Language Audio & Mus, London, England
来源
关键词
affective computing; negative interactions; cross-corpora analysis; conflict escalation; speech paralinguistics; emotion recognition; ACOUSTIC EMOTION RECOGNITION; CONFLICT;
D O I
10.3389/fcomp.2022.749804
中图分类号
TP39 [计算机的应用];
学科分类号
081203 ; 0835 ;
摘要
The monitoring of an escalating negative interaction has several benefits, particularly in security, (mental) health, and group management. The speech signal is particularly suited to this, as aspects of escalation, including emotional arousal, are proven to easily be captured by the audio signal. A challenge of applying trained systems in real-life applications is their strong dependence on the training material and limited generalization abilities. For this reason, in this contribution, we perform an extensive analysis of three corpora in the Dutch language. All three corpora are high in escalation behavior content and are annotated on alternative dimensions related to escalation. A process of label mapping resulted in two possible ground truth estimations for the three datasets as low, medium, and high escalation levels. To observe class behavior and inter-corpus differences more closely, we perform acoustic analysis of the audio samples, finding that derived labels perform similarly across each corpus, with escalation interaction increasing in pitch (F0) and intensity (dB). We explore the suitability of different speech features, data augmentation, merging corpora for training, and testing on actor and non-actor speech through our experiments. We find that the extent to which merging corpora is successful depends greatly on the similarities between label definitions before label mapping. Finally, we see that the escalation recognition task can be performed in a cross-corpus setup with hand-crafted speech features, obtaining up to 63.8% unweighted average recall (UAR) at best for a cross-corpus analysis, an increase from the inter-corpus results of 59.4% UAR.
引用
收藏
页数:13
相关论文
共 50 条
  • [31] Cross-corpus Speech Emotion Recognition Using Transfer Semi-supervised Discriminant Analysis
    Song, Peng
    Zhang, Xinran
    Ou, Shifeng
    Liu, Jingjing
    Yu, Yanwei
    Zheng, Wenming
    2016 10TH INTERNATIONAL SYMPOSIUM ON CHINESE SPOKEN LANGUAGE PROCESSING (ISCSLP), 2016,
  • [32] Nonnegative Matrix Factorization Based Transfer Subspace Learning for Cross-Corpus Speech Emotion Recognition
    Luo, Hui
    Han, Jiqing
    IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2020, 28 : 2047 - 2060
  • [33] Exploring corpus-invariant emotional acoustic feature for cross-corpus speech emotion recognition
    Lian, Hailun
    Lu, Cheng
    Zhao, Yan
    Li, Sunan
    Qi, Tianhua
    Zong, Yuan
    EXPERT SYSTEMS WITH APPLICATIONS, 2024, 258
  • [34] Deep Transductive Transfer Regression Network for Cross-Corpus Speech Emotion Recognition
    Zhao, Yan
    Wang, Jincen
    Ye, Ru
    Zong, Yuan
    Zheng, Wenming
    Zhao, Li
    INTERSPEECH 2022, 2022, : 371 - 375
  • [35] Towards Domain-Specific Cross-Corpus Speech Emotion Recognition Approach
    Zhao, Yan
    Zong, Yuan
    Lian, Hailun
    Lu, Cheng
    Shi, Jingang
    Zheng, Wenming
    IEEE TRANSACTIONS ON COMPUTATIONAL SOCIAL SYSTEMS, 2024,
  • [36] Learning Complex Spectral Mapping for Speech Enhancement with Improved Cross-corpus Generalization
    Pandey, Ashutosh
    Wang, DeLiang
    INTERSPEECH 2020, 2020, : 4511 - 4515
  • [37] Target-Adapted Subspace Learning for Cross-Corpus Speech Emotion Recognition
    Chen, Xiuzhen
    Zhou, Xiaoyan
    Lu, Cheng
    Zong, Yuan
    Zheng, Wenming
    Tang, Chuangao
    IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2019, E102D (12) : 2632 - 2636
  • [38] CROSS-CORPUS EEG-BASED EMOTION RECOGNITION
    Rayatdoost, Soheil
    Soleymani, Mohammad
    2018 IEEE 28TH INTERNATIONAL WORKSHOP ON MACHINE LEARNING FOR SIGNAL PROCESSING (MLSP), 2018,
  • [39] An adaptation framework with unified embedding reconstruction for cross-corpus speech emotion recognition
    Zhang, Ruiteng
    Wei, Jianguo
    Lu, Xugang
    Li, Yongwei
    Lu, Wenhuan
    Zhang, Lin
    Xu, Junhai
    APPLIED SOFT COMPUTING, 2025, 174
  • [40] Within and cross-corpus speech emotion recognition using latent topic model-based features
    Mohit Shah
    Chaitali Chakrabarti
    Andreas Spanias
    EURASIP Journal on Audio, Speech, and Music Processing, 2015