Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms

被引:31
|
作者
Tursunov, Anvarjon [1 ]
Mustageem [1 ]
Choeh, Joon Yeon [2 ]
Kwon, Soonil [1 ]
机构
[1] Sejong Univ, Dept Software, Interact Technol Lab, Seoul 05006, South Korea
[2] Sejong Univ, Dept Software, Intelligent Contents Lab, Seoul 05006, South Korea
关键词
human-computer interaction; convolutional neural network; multi-attention module; age and gender recognition; speech signals; SPEAKER AGE; DEEP; CLASSIFICATION; LSTM; CNN;
D O I
10.3390/s21175892
中图分类号
O65 [分析化学];
学科分类号
070302 ; 081704 ;
摘要
Speech signals are being used as a primary input source in human-computer interaction (HCI) to develop several applications, such as automatic speech recognition (ASR), speech emotion recognition (SER), gender, and age recognition. Classifying speakers according to their age and gender is a challenging task in speech processing owing to the disability of the current methods of extracting salient high-level speech features and classification models. To address these problems, we introduce a novel end-to-end age and gender recognition convolutional neural network (CNN) with a specially designed multi-attention module (MAM) from speech signals. Our proposed model uses MAM to extract spatial and temporal salient features from the input data effectively. The MAM mechanism uses a rectangular shape filter as a kernel in convolution layers and comprises two separate time and frequency attention mechanisms. The time attention branch learns to detect temporal cues, whereas the frequency attention module extracts the most relevant features to the target by focusing on the spatial frequency features. The combination of the two extracted spatial and temporal features complements one another and provide high performance in terms of age and gender classification. The proposed age and gender classification system was tested using the Common Voice and locally developed Korean speech recognition datasets. Our suggested model achieved 96%, 73%, and 76% accuracy scores for gender, age, and age-gender classification, respectively, using the Common Voice dataset. The Korean speech recognition dataset results were 97%, 97%, and 90% for gender, age, and age-gender recognition, respectively. The prediction performance of our proposed model, which was obtained in the experiments, demonstrated the superiority and robustness of the tasks regarding age, gender, and age-gender recognition from speech signals.
引用
收藏
页数:19
相关论文
共 50 条
  • [41] Gender Classification in Speech Recognition using Fuzzy Logic and Neural Network
    Meena, Kunjithapatham
    Subramaniam, Kulumani
    Gomathy, Muthusamy
    INTERNATIONAL ARAB JOURNAL OF INFORMATION TECHNOLOGY, 2013, 10 (05) : 477 - 485
  • [42] Speckle Noise Removal in Ultrasound Images Using a Deep Convolutional Neural Network and a Specially Designed Loss Function
    Feng, Danlei
    Wu, Weichen
    Li, Hongfeng
    Li, Quanzheng
    MULTISCALE MULTIMODAL MEDICAL IMAGING, MMMI 2019, 2020, 11977 : 85 - 92
  • [43] Attention induced multi-head convolutional neural network for human activity recognition
    Khan, Zanobya N.
    Ahmad, Jamil
    APPLIED SOFT COMPUTING, 2021, 110
  • [44] Multimodal speech emotion recognition and classification using convolutional neural network techniques
    A. Christy
    S. Vaithyasubramanian
    A. Jesudoss
    M. D. Anto Praveena
    International Journal of Speech Technology, 2020, 23 : 381 - 388
  • [45] Multimodal speech emotion recognition and classification using convolutional neural network techniques
    Christy, A.
    Vaithyasubramanian, S.
    Jesudoss, A.
    Praveena, M. D. Anto
    INTERNATIONAL JOURNAL OF SPEECH TECHNOLOGY, 2020, 23 (02) : 381 - 388
  • [46] Convolutional Neural Network applied in mime speech recognition using sEMG data
    Ai, Qing
    Zhang, Wei
    Zhang, Bixuan
    Li, Guang
    Yang, Meng
    2019 CHINESE AUTOMATION CONGRESS (CAC2019), 2019, : 3347 - 3352
  • [47] Developing a Speech Recognition System for Recognizing Tonal Speech Signals Using a Convolutional Neural Network
    Dua, Sakshi
    Kumar, Sethuraman Sambath
    Albagory, Yasser
    Ramalingam, Rajakumar
    Dumka, Ankur
    Singh, Rajesh
    Rashid, Mamoon
    Gehlot, Anita
    Alshamrani, Sultan S.
    AlGhamdi, Ahmed Saeed
    APPLIED SCIENCES-BASEL, 2022, 12 (12):
  • [48] Multi-attention Deep Recurrent Neural Network for Nursing Action Evaluation Using Wearable Sensor
    Zhong, Zhihang
    Lin, Chingszu
    Ogata, Taiki
    Ota, Jun
    PROCEEDINGS OF THE 25TH INTERNATIONAL CONFERENCE ON INTELLIGENT USER INTERFACES, IUI 2020, 2020, : 546 - 550
  • [49] A Multi-View Gait Recognition Method Using Deep Convolutional Neural Network and Channel Attention Mechanism
    Wang, Jiabin
    Peng, Kai
    CMES-COMPUTER MODELING IN ENGINEERING & SCIENCES, 2020, 125 (01): : 345 - 363
  • [50] Attention-Based Multi-Filter Convolutional Neural Network for Inappropriate Speech Detection
    Lin, Shu-Yu
    Chen, Yi-Ling
    2021 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS (IJCNN), 2021,