Neural Spatio-Temporal Beamformer for Target Speech Separation

被引:20
|
作者
Xu, Yong [1 ]
Yu, Meng [1 ]
Zhang, Shi-Xiong [1 ]
Chen, Lianwu [2 ]
Weng, Chao [1 ]
Liu, Jianming [1 ]
Yu, Dong [1 ]
机构
[1] Tencent AI Lab, Bellevue, WA 98004 USA
[2] Tencent AI Lab, Shenzhen, Peoples R China
来源
关键词
target speech separation; multi-tap MVDR; mask-based MVDR; spatio-temporal beamformer; NOISE-REDUCTION; ENHANCEMENT; RECOGNITION; END;
D O I
10.21437/Interspeech.2020-1458
中图分类号
R36 [病理学]; R76 [耳鼻咽喉科学];
学科分类号
100104 ; 100213 ;
摘要
Purely neural network (NN) based speech separation and enhancement methods, although can achieve good objective scores, inevitably cause nonlinear speech distortions that are harmful for the automatic speech recognition (ASR). On the other hand, the minimum variance distortionless response (MVDR) beamformer with NN-predicted masks, although can significantly reduce speech distortions, has limited noise reduction capability. In this paper, we propose a multi-tap MVDR beamformer with complex-valued masks for speech separation and enhancement. Compared to the state-of-the-art NN-mask based MVDR beamformer, the multi-tap MVDR beamformer exploits the inter-frame correlation in addition to the inter-microphone correlation that is already utilized in prior arts. Further improvements include the replacement of the real-valued masks with the complex-valued masks and the joint training of the complex-mask NN. The evaluation on our multi-modal multi-channel target speech separation and enhancement platform demonstrates that our proposed multi-tap MVDR beamformer improves both the ASR accuracy and the perceptual speech quality against prior arts.
引用
收藏
页码:56 / 60
页数:5
相关论文
共 50 条
  • [31] Improved Target Tracking Based on Spatio-Temporal Learning
    Jia, Songmin
    Zeng, Dishi
    Xu, Tao
    Zhang, Hui
    Li, Xiuzhi
    2016 IEEE INTERNATIONAL CONFERENCE ON INFORMATION AND AUTOMATION (ICIA), 2016, : 1840 - 1845
  • [32] A spatio-temporal speech enhancement scheme for robust speech recognition in noisy environments
    Visser, E
    Otsuka, M
    Lee, TW
    SPEECH COMMUNICATION, 2003, 41 (2-3) : 393 - 407
  • [33] Spatio-Temporal Prediction of Suspect Location by Spatio-Temporal Semantics
    Duan L.
    Hu T.
    Zhu X.
    Ye X.
    Wang S.
    Wuhan Daxue Xuebao (Xinxi Kexue Ban)/Geomatics and Information Science of Wuhan University, 2019, 44 (05): : 765 - 770
  • [34] Habituation based neural networks for spatio-temporal classification
    Stiles, BW
    Ghosh, J
    NEUROCOMPUTING, 1997, 15 (3-4) : 273 - 307
  • [35] Spatio-temporal image filtering with cellular neural networks
    Shi, BE
    ICNN - 1996 IEEE INTERNATIONAL CONFERENCE ON NEURAL NETWORKS, VOLS. 1-4, 1996, : 1410 - 1415
  • [36] Exploring the spatio-temporal neural basis of face learning
    Yang, Ying
    Xu, Yang
    Jew, Carol A.
    Pyles, John A.
    Kass, Robert E.
    Tarr, Michael J.
    JOURNAL OF VISION, 2017, 17 (06):
  • [37] On the inclusion of spatial information for spatio-temporal neural networks
    de Medrano, Rodrigo
    Aznarte, Jose L.
    NEURAL COMPUTING & APPLICATIONS, 2021, 33 (21): : 14723 - 14740
  • [38] Misbehavior detection with spatio-temporal graph neural networks ☆
    Yuce, Mehmet Fatih
    Erturk, Mehmet Ali
    Aydin, Muhammed Ali
    COMPUTERS & ELECTRICAL ENGINEERING, 2024, 116
  • [39] Spatio-temporal sequence processing with the counterpropagation neural network
    Wang, LP
    INFORMATION INTELLIGENCE AND SYSTEMS, VOLS 1-4, 1996, : 831 - 835
  • [40] Spatio-temporal Representations of Uncertainty in Spiking Neural Networks
    Savin, Cristina
    Deneve, Sophie
    ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 27 (NIPS 2014), 2014, 27