Age and Gender Recognition Using a Convolutional Neural Network with a Specially Designed Multi-Attention Module through Speech Spectrograms

被引:31
|
作者
Tursunov, Anvarjon [1 ]
Mustageem [1 ]
Choeh, Joon Yeon [2 ]
Kwon, Soonil [1 ]
机构
[1] Sejong Univ, Dept Software, Interact Technol Lab, Seoul 05006, South Korea
[2] Sejong Univ, Dept Software, Intelligent Contents Lab, Seoul 05006, South Korea
关键词
human-computer interaction; convolutional neural network; multi-attention module; age and gender recognition; speech signals; SPEAKER AGE; DEEP; CLASSIFICATION; LSTM; CNN;
D O I
10.3390/s21175892
中图分类号
O65 [分析化学];
学科分类号
070302 ; 081704 ;
摘要
Speech signals are being used as a primary input source in human-computer interaction (HCI) to develop several applications, such as automatic speech recognition (ASR), speech emotion recognition (SER), gender, and age recognition. Classifying speakers according to their age and gender is a challenging task in speech processing owing to the disability of the current methods of extracting salient high-level speech features and classification models. To address these problems, we introduce a novel end-to-end age and gender recognition convolutional neural network (CNN) with a specially designed multi-attention module (MAM) from speech signals. Our proposed model uses MAM to extract spatial and temporal salient features from the input data effectively. The MAM mechanism uses a rectangular shape filter as a kernel in convolution layers and comprises two separate time and frequency attention mechanisms. The time attention branch learns to detect temporal cues, whereas the frequency attention module extracts the most relevant features to the target by focusing on the spatial frequency features. The combination of the two extracted spatial and temporal features complements one another and provide high performance in terms of age and gender classification. The proposed age and gender classification system was tested using the Common Voice and locally developed Korean speech recognition datasets. Our suggested model achieved 96%, 73%, and 76% accuracy scores for gender, age, and age-gender classification, respectively, using the Common Voice dataset. The Korean speech recognition dataset results were 97%, 97%, and 90% for gender, age, and age-gender recognition, respectively. The prediction performance of our proposed model, which was obtained in the experiments, demonstrated the superiority and robustness of the tasks regarding age, gender, and age-gender recognition from speech signals.
引用
收藏
页数:19
相关论文
共 50 条
  • [1] Electromyographic hand gesture recognition using convolutional neural network with multi-attention
    Zhang, Zhen
    Shen, Quming
    Wang, Yanyu
    [J]. BIOMEDICAL SIGNAL PROCESSING AND CONTROL, 2024, 91
  • [2] Multi-Attention Convolutional Neural Network for Video Deblurring
    Zhang, Xiaoqin
    Wang, Tao
    Jiang, Runhua
    Zhao, Li
    Xu, Yuewang
    [J]. IEEE TRANSACTIONS ON CIRCUITS AND SYSTEMS FOR VIDEO TECHNOLOGY, 2022, 32 (04) : 1986 - 1997
  • [3] Adaptive Multi-Attention Convolutional Neural Network for Fine-Grained Image Recognition
    Li, Ang
    Chen, Jianxin
    Kang, Bin
    Zhuang, Wenqin
    Zhang, Xuguang
    [J]. 2019 IEEE GLOBECOM WORKSHOPS (GC WKSHPS), 2019,
  • [4] Learning Multi-Attention Convolutional Neural Network for Fine-Grained Image Recognition
    Zheng, Heliang
    Fu, Jianlong
    Mei, Tao
    Luo, Jiebo
    [J]. 2017 IEEE INTERNATIONAL CONFERENCE ON COMPUTER VISION (ICCV), 2017, : 5219 - 5227
  • [5] Speech Emotion Recognition from Spectrograms with Deep Convolutional Neural Network
    Badshah, Abdul Malik
    Ahmad, Jamil
    Rahim, Nasir
    Baik, Sung Wook
    [J]. 2017 INTERNATIONAL CONFERENCE ON PLATFORM TECHNOLOGY AND SERVICE (PLATCON), 2017, : 125 - 129
  • [6] Speech Emotion Recognition using Convolutional Recurrent Neural Networks and Spectrograms
    Qamhan, Mustafa A.
    Meftah, Ali H.
    Selouani, Sid-Ahmed
    Alotaibi, Yousef A.
    Zakariah, Mohammed
    Seddiq, Yasser Mohammad
    [J]. 2020 IEEE CANADIAN CONFERENCE ON ELECTRICAL AND COMPUTER ENGINEERING (CCECE), 2020,
  • [7] Application of convolutional neural network for gender and age group recognition from speech
    Pham Tuan Dat
    Le The Anh
    [J]. PROCEEDINGS OF 2019 6TH NATIONAL FOUNDATION FOR SCIENCE AND TECHNOLOGY DEVELOPMENT (NAFOSTED) CONFERENCE ON INFORMATION AND COMPUTER SCIENCE (NICS), 2019, : 489 - 493
  • [8] Bearing Fault Diagnosis Using Convolutional Neural Network Based on a Multi-Attention Mechanism
    Kang T.
    Duan R.
    Yang L.
    Xue J.
    Liao Y.
    [J]. Hsi-An Chiao Tung Ta Hsueh/Journal of Xi'an Jiaotong University, 2022, 56 (12): : 68 - 77
  • [9] Shearlet Convolutional Neural Network Approach for Age and Gender recognition
    Ziani, Chaymae
    Sadiq, Abdelalim
    [J]. 2019 THIRD INTERNATIONAL CONFERENCE ON INTELLIGENT COMPUTING IN DATA SCIENCES (ICDS 2019), 2019,
  • [10] Spatiotemporal Convolutional Neural Network with Convolutional Block Attention Module for Micro-Expression Recognition
    Chen, Boyu
    Zhang, Zhihao
    Liu, Nian
    Tan, Yang
    Liu, Xinyu
    Chen, Tong
    [J]. INFORMATION, 2020, 11 (08)