Optimizations of Neural Audio Coder Toward Perceptual Transparency

被引:0
|
作者
Byun, Joon [1 ]
Shin, Seungmin [1 ]
Hwang, Seorim [1 ]
Sung, Jongmo [2 ]
Beack, Seungkwon [2 ]
Park, Youngcheol [3 ]
机构
[1] Yonsei Univ, Comp Sci Dept, Wonju 26493, South Korea
[2] Elect & Telecommun Res Inst, Daejeon 34129, South Korea
[3] Yonsei Univ, Software Div, Wonju 26493, South Korea
关键词
Psychoacoustic models; Entropy; Encoding; Distortion measurement; Optimization; Distortion; Noise; Neural audio coder; optimization; psychoacoustic model; perceptual loss function; entropy model; hyperprior;
D O I
10.1109/JSTSP.2024.3437155
中图分类号
TM [电工技术]; TN [电子技术、通信技术];
学科分类号
0808 ; 0809 ;
摘要
This paper presents comprehensive optimizations of a neural audio coder built upon a variational autoencoder (VAE) system integrated with an arithmetic coder. Our optimizations focus on two primary aspects: a novel loss function design and advanced entropy modeling of bottleneck latent embeddings. The loss function design incorporates parameters from a psychoacoustic model (PAM) into the frame-wise distortion measure, providing excellent perceptual quality. In addition, a multi-time scale discriminator is utilized to minimize distortions across adjacent frames, reducing artifacts at frame edges. Also, the coder is optimized considering three sophisticated entropy models within the latent domain: the Factorized Entropy Model (FEM), the Hyperprior Model (HPM), and the Joint Hierarchical Model (JHM). Notably, the JHM enhances context modeling across frames to effectively predict components influenced by long-term dependencies. To verify the optimization performance, we conducted extensive experiments using a dataset consisting of commercial movie clips and two additional public datasets. Objective metrics consistently demonstrated that our optimized loss function and latent modeling achieved superior performance across all test datasets compared to traditional codecs such as LAME-MP3 and FDK-AAC. Subjective assessments also indicated that our system could offer comparable or superior auditory quality to FDK-AAC.
引用
收藏
页码:1531 / 1543
页数:13
相关论文
共 50 条
  • [21] Perceptual asymmetry reveals neural substrates underlying stereoscopic transparency
    Tsirlin, Inna
    Allison, Robert S.
    Wilcox, Laurie M.
    VISION RESEARCH, 2012, 54 : 1 - 11
  • [22] DEEP NEURAL NETWORK (DNN) AUDIO CODER USING A PERCEPTUALLY IMPROVED TRAINING METHOD
    Shin, Seungmin
    Byun, Joon
    Park, Youngcheol
    Sung, Jongmo
    Beack, Seungkwon
    2022 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), 2022, : 871 - 875
  • [23] Perceptual Neural Audio Coding With Modified Discrete Cosine Transform
    Lim, Hyungseob
    Lee, Jihyun
    Kim, Byeong Hyeon
    Jang, Inseon
    Kang, Hong-Goo
    IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, 2024, 18 (08) : 1490 - 1505
  • [24] A switched parametric & transform audio coder
    Levine, SN
    Smith, JO
    ICASSP '99: 1999 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH, AND SIGNAL PROCESSING, PROCEEDINGS VOLS I-VI, 1999, : 985 - 988
  • [25] Switched parametric & transform audio coder
    Levine, Scott N.
    Smith III, Julius O.
    ICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings, 1999, 2 : 985 - 988
  • [26] Low complexity scalable perceptual audio coder using an optimum wavelet packet basis representation and vector quantization
    Sathidevi, PS
    Venkataramani, Y
    IETE JOURNAL OF RESEARCH, 2004, 50 (06) : 399 - 407
  • [27] Neural-Based Approach to Perceptual Sparse Coding of Audio Signals
    Pichevar, Ramin
    Najaf-Zadeh, Hossein
    Mustiere, Frederic
    2010 INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS IJCNN 2010, 2010,
  • [28] The Limitations of Perceptual Transparency
    Gow, Laura
    PHILOSOPHICAL QUARTERLY, 2016, 66 (265): : 723 - 744
  • [29] Conditions for perceptual transparency
    Ripamonti, C
    Westland, S
    Da Pos, O
    JOURNAL OF ELECTRONIC IMAGING, 2004, 13 (01) : 29 - 35
  • [30] Conditions for perceptual transparency
    Westland, S
    Da Pos, O
    Ripamonti, C
    HUMAN VISION AND ELECTRONIC IMAGING VII, 2002, 4662 : 315 - 323