ROBUST TEXT-INDEPENDENT SPEAKER IDENTIFICATION USING GAUSSIAN MIXTURE SPEAKER MODELS

被引:1721
|
作者
REYNOLDS, DA
ROSE, RC
机构
[1] Speech Systems Technology Group, MIT Lincoln Laboratory, Lexington
[2] Speech Research Department, AT&T Bell Laboratories, Murray Hill
来源
关键词
D O I
10.1109/89.365379
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
This paper introduces and motivates the use of Gaussian mixture models (GMM) for robust text-independent speaker identification, The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity, The focus of this work is on applications which require high identification rates using short utterance from unconstrained conversational speech and robustness to degradations produced by transmission over a telephone channel, A complete experimental evaluation of the Gaussian mixture speaker model is conducted on a 49 speaker, conversational telephone speech database, The experiments examine algorithmic issues (initialization, variance limiting, model order selection), spectral variability robustness techniques, large population performance, and comparisons to other speaker modeling techniques (uni-modal Gaussian, VQ codebook, tied Gaussian mixture, and radial basis functions), The Gaussian mixture speaker model attains 96.8% identification accuracy using 5 second clean speech utterances and 80.8% accuracy using 15 second telephone speech utterances with a 49 speaker population and is shown to outperform the other speaker modeling techniques on an identical 16 speaker telephone speech task.
引用
收藏
页码:72 / 83
页数:12
相关论文
共 50 条
  • [1] Improved Text-Independent Speaker Identification and Verification with Gaussian Mixture Models
    Chakroun, Rania
    Frikha, Mondher
    [J]. KNOWLEDGE SCIENCE, ENGINEERING AND MANAGEMENT, KSEM 2019, PT II, 2019, 11776 : 3 - 10
  • [2] Robust Text-independent Speaker recognition with Short Utterances using Gaussian Mixture Models
    Chakroun, Rania
    Frikha, Mondher
    [J]. 2020 16TH INTERNATIONAL WIRELESS COMMUNICATIONS & MOBILE COMPUTING CONFERENCE, IWCMC, 2020, : 2204 - 2209
  • [3] Frame level likelihood normalization for text-independent speaker identification using Gaussian Mixture Models
    Markov, K
    Nakagawa, S
    [J]. ICSLP 96 - FOURTH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, PROCEEDINGS, VOLS 1-4, 1996, : 1764 - 1767
  • [4] Dimensionality reduction for text-independent speaker identification using Gaussian Mixture Model
    El-Gamal, MA
    Abu El-Yazeed, MF
    El Ayadi, MMH
    [J]. Proceedings of the 46th IEEE International Midwest Symposium on Circuits & Systems, Vols 1-3, 2003, : 625 - 628
  • [5] Text-Independent Speaker Verification Using Variational Gaussian Mixture Model
    Moattar, Mohammad Hossein
    Homayounpour, Mohammad Mehdi
    [J]. ETRI JOURNAL, 2011, 33 (06) : 914 - 923
  • [6] Text-independent speaker identification using robust statistics estimation
    El Ayadi, Moataz
    Hassan, Abdel-Karim S. O.
    Abdel-Naby, Ahmed
    Elgendy, Omar A.
    [J]. SPEECH COMMUNICATION, 2017, 92 : 52 - 63
  • [7] Robust text-independent speaker identification using bispectrum slice
    Özkurt, TE
    Akgül, T
    [J]. PROCEEDINGS OF THE IEEE 12TH SIGNAL PROCESSING AND COMMUNICATIONS APPLICATIONS CONFERENCE, 2004, : 418 - 421
  • [8] Self-Organizing Mixture Models for Text-Independent Speaker Identification
    Bouziane, Ayoub
    Kharroubi, Jamal
    Zarghili, Arsalane
    [J]. 2014 THIRD IEEE INTERNATIONAL COLLOQUIUM IN INFORMATION SCIENCE AND TECHNOLOGY (CIST'14), 2014, : 345 - 350
  • [9] Text-independent speaker identification using Gaussian mixture models based on multi-space probability distribution
    Miyajima, C
    Hattori, Y
    Tokuda, K
    Masuko, T
    Kobayashi, T
    Kitamura, T
    [J]. IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2001, E84D (07): : 847 - 855
  • [10] A Chain of Gaussian Mixture Model for Text-independent Speaker Recognition
    Chen, Yanxiang
    Liu, Ming
    [J]. ORIENTAL COCOSDA 2009 - INTERNATIONAL CONFERENCE ON SPEECH DATABASE AND ASSESSMENTS, 2009, : 100 - +