Statistical Model-Based Voice Activity Detection Using Spatial Cues and Log Energy for Dual-Channel Noisy Speech Recognition

被引:0
|
作者
Park, Ji Hun [1 ]
Shin, Min Hwa [2 ]
Kim, Hong Kook [1 ]
机构
[1] Gwangju Inst Sci & Technol, Sch Informat & Commun, Kwangju 500712, South Korea
[2] Multimedia IP Res Ctr, Korea Elect Technol Inst, Seongnam 463816, South Korea
来源
基金
新加坡国家研究基金会;
关键词
Voice activity detection (VAD); end-point detection; dual-channel speech; speech recognition; spatial cues;
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
In this paper, a voice activity detection (VAD) method for dual-channel noisy speech recognition is proposed on the basis of statistical models constructed by spatial cues and log energy. In particular, spatial cues are composed of the interaural time differences and interaural level differences of dual-channel speech signals, and the statistical models for speech presence and absence are based on a Gaussian kernel density. In order to evaluate the performance of the proposed VAD method, speech recognition is performed using only speech signals segmented by the proposed VAD method. The performance of the proposed VAD method is then compared with those of conventional methods such as a signal-to-noise ratio variance based method and a phase vector based method. It is shown from the experiments that the proposed VAD method outperforms conventional methods, providing the relative word error rate reductions of 19.5% and 12.2%, respectively.
引用
收藏
页码:172 / +
页数:2
相关论文
共 42 条
  • [1] Speech recognition enhancement with statistical model-based voice activity detection
    Jarc, Bojan
    Babič, Rudolf
    Elektrotehniski Vestnik/Electrotechnical Review, 2002, 69 (01): : 75 - 81
  • [2] A statistical model-based voice activity detection
    Sohn, J
    Kim, NS
    Sung, W
    IEEE SIGNAL PROCESSING LETTERS, 1999, 6 (01) : 1 - 3
  • [3] Dual Microphone Voice Activity Detection Based on Reliable Spatial Cues
    Hwang, Soojoong
    Jin, Yu Gwang
    Shin, Jong Won
    SENSORS, 2019, 19 (14)
  • [4] Voice Activity Detection Method Using Psycho Acoustic Model Based on Speech Energy Maximization in Noisy Environments
    Choi, Gab-Keun
    Kim, Soon-Hyob
    JOURNAL OF THE ACOUSTICAL SOCIETY OF KOREA, 2009, 28 (05): : 447 - 453
  • [5] Statistical model-based voice activity detection using support vector machine
    Jo, Q-H.
    Chang, J. -H.
    Shin, J. W.
    Kim, N. S.
    IET SIGNAL PROCESSING, 2009, 3 (03) : 205 - 210
  • [6] A novel voice activity detection based on phoneme recognition using statistical model
    Bao, Xulei
    Zhu, Jie
    EURASIP JOURNAL ON AUDIO SPEECH AND MUSIC PROCESSING, 2012,
  • [7] A novel voice activity detection based on phoneme recognition using statistical model
    Xulei Bao
    Jie Zhu
    EURASIP Journal on Audio, Speech, and Music Processing, 2012
  • [8] Postfilter for Dual Channel Speech Enhancement Using Coherence and Statistical Model-Based Noise Estimation
    Cheong, Sein
    Kim, Minseung
    Shin, Jong Won
    SENSORS, 2024, 24 (12)
  • [9] A Statistical Model-Based Voice Activity Detection Using Multiple DNNs and Noise Awareness
    Hwang, Inyoung
    Sim, Jaeseong
    Kim, Sang-Hyeon
    Song, Kwang-Sub
    Chang, Joon-Hyuk
    16TH ANNUAL CONFERENCE OF THE INTERNATIONAL SPEECH COMMUNICATION ASSOCIATION (INTERSPEECH 2015), VOLS 1-5, 2015, : 2277 - 2281
  • [10] Discriminative Weight Training for a Statistical Model-Based Voice Activity Detection
    Kang, Sang-Ick
    Jo, Q-Haing
    Chang, Joon-Hyuk
    Park, Seung Seop
    JOURNAL OF THE ACOUSTICAL SOCIETY OF KOREA, 2007, 26 (05): : 194 - 198