ROBUST TEXT-INDEPENDENT SPEAKER IDENTIFICATION USING GAUSSIAN MIXTURE SPEAKER MODELS

被引:1721
|
作者
REYNOLDS, DA
ROSE, RC
机构
[1] Speech Systems Technology Group, MIT Lincoln Laboratory, Lexington
[2] Speech Research Department, AT&T Bell Laboratories, Murray Hill
来源
关键词
D O I
10.1109/89.365379
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
This paper introduces and motivates the use of Gaussian mixture models (GMM) for robust text-independent speaker identification, The individual Gaussian components of a GMM are shown to represent some general speaker-dependent spectral shapes that are effective for modeling speaker identity, The focus of this work is on applications which require high identification rates using short utterance from unconstrained conversational speech and robustness to degradations produced by transmission over a telephone channel, A complete experimental evaluation of the Gaussian mixture speaker model is conducted on a 49 speaker, conversational telephone speech database, The experiments examine algorithmic issues (initialization, variance limiting, model order selection), spectral variability robustness techniques, large population performance, and comparisons to other speaker modeling techniques (uni-modal Gaussian, VQ codebook, tied Gaussian mixture, and radial basis functions), The Gaussian mixture speaker model attains 96.8% identification accuracy using 5 second clean speech utterances and 80.8% accuracy using 15 second telephone speech utterances with a 49 speaker population and is shown to outperform the other speaker modeling techniques on an identical 16 speaker telephone speech task.
引用
收藏
页码:72 / 83
页数:12
相关论文
共 50 条
  • [41] A robust DNN model for text-independent speaker identification using non-speaker embeddings in diverse data conditions
    Nirupam Shome
    Banala Saritha
    Richik Kashyap
    Rabul Hussain Laskar
    [J]. Neural Computing and Applications, 2023, 35 : 18933 - 18947
  • [42] A robust DNN model for text-independent speaker identification using non-speaker embeddings in diverse data conditions
    Shome, Nirupam
    Saritha, Banala
    Kashyap, Richik
    Laskar, Rabul Hussain
    [J]. NEURAL COMPUTING & APPLICATIONS, 2023, 35 (26): : 18933 - 18947
  • [43] A robust sequential test for text-independent speaker verification
    Lund, MA
    Lee, CC
    [J]. JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1996, 99 (01): : 609 - 621
  • [44] On Von-Mises Fisher Mixture Model in Text-Independent Speaker Identification
    Taghia, Jalil
    Ma, Zhanyu
    Leijon, Arne
    [J]. 14TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION (INTERSPEECH 2013), VOLS 1-5, 2013, : 2498 - 2502
  • [45] TEXT-INDEPENDENT SPEAKER RECOGNITION
    ATAL, BS
    [J]. JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA, 1972, 52 (01): : 181 - &
  • [46] TEXT-INDEPENDENT SPEAKER IDENTIFICATION USING NEURAL NETS AND AR-VECTOR MODELS
    HADJITODOROV, S
    BOYANOV, B
    IVANOV, T
    DALAKCHIEVA, N
    [J]. ELECTRONICS LETTERS, 1994, 30 (11) : 838 - 840
  • [47] A Novel Approach in Feature Level for Robust Text-Independent Speaker Identification system
    Sarangi, Susanta Kumar
    Saha, Goutam
    [J]. 4TH INTERNATIONAL CONFERENCE ON INTELLIGENT HUMAN COMPUTER INTERACTION (IHCI 2012), 2012,
  • [48] A New Set of Features for Text-Independent Speaker Identification
    Espy-Wilson, Carol Y.
    Manocha, Sandeep
    Vishnubhotla, Srikanth
    [J]. INTERSPEECH 2006 AND 9TH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, VOLS 1-5, 2006, : 1475 - +
  • [49] Research on noise compensation for text-independent speaker identification
    Qiu, Hong
    Wu, Shuzhen
    [J]. Beijing Daxue Xuebao (Ziran Kexue Ban)/Acta Scientiarum Naturalium Universitatis Pekinensis, 2005, 41 (01): : 115 - 121
  • [50] A Novel Reduction Method for Text-Independent Speaker Identification
    Wang, Yan
    Liu, Xueyan
    Xing, Yujuan
    Li, Ming
    [J]. ICNC 2008: FOURTH INTERNATIONAL CONFERENCE ON NATURAL COMPUTATION, VOL 4, PROCEEDINGS, 2008, : 66 - 70