Optimizations of Neural Audio Coder Toward Perceptual Transparency

被引:0
|
作者
Byun, Joon [1 ]
Shin, Seungmin [1 ]
Hwang, Seorim [1 ]
Sung, Jongmo [2 ]
Beack, Seungkwon [2 ]
Park, Youngcheol [3 ]
机构
[1] Yonsei Univ, Comp Sci Dept, Wonju 26493, South Korea
[2] Elect & Telecommun Res Inst, Daejeon 34129, South Korea
[3] Yonsei Univ, Software Div, Wonju 26493, South Korea
关键词
Psychoacoustic models; Entropy; Encoding; Distortion measurement; Optimization; Distortion; Noise; Neural audio coder; optimization; psychoacoustic model; perceptual loss function; entropy model; hyperprior;
D O I
10.1109/JSTSP.2024.3437155
中图分类号
TM [电工技术]; TN [电子技术、通信技术];
学科分类号
0808 ; 0809 ;
摘要
This paper presents comprehensive optimizations of a neural audio coder built upon a variational autoencoder (VAE) system integrated with an arithmetic coder. Our optimizations focus on two primary aspects: a novel loss function design and advanced entropy modeling of bottleneck latent embeddings. The loss function design incorporates parameters from a psychoacoustic model (PAM) into the frame-wise distortion measure, providing excellent perceptual quality. In addition, a multi-time scale discriminator is utilized to minimize distortions across adjacent frames, reducing artifacts at frame edges. Also, the coder is optimized considering three sophisticated entropy models within the latent domain: the Factorized Entropy Model (FEM), the Hyperprior Model (HPM), and the Joint Hierarchical Model (JHM). Notably, the JHM enhances context modeling across frames to effectively predict components influenced by long-term dependencies. To verify the optimization performance, we conducted extensive experiments using a dataset consisting of commercial movie clips and two additional public datasets. Objective metrics consistently demonstrated that our optimized loss function and latent modeling achieved superior performance across all test datasets compared to traditional codecs such as LAME-MP3 and FDK-AAC. Subjective assessments also indicated that our system could offer comparable or superior auditory quality to FDK-AAC.
引用
收藏
页码:1531 / 1543
页数:13
相关论文
共 50 条
  • [1] Perceptual Audio Quality Assessment for Coder Evaluation
    Garcia-Alvarez, Julio C.
    Aguirre, Santiago E.
    Diaz-Solarte, Paulo C.
    2014 IEEE Fourth International Conference on Consumer Electronics Berlin (ICCE-Berlin), 2014, : 408 - 410
  • [2] On the use of backward adaptation in a perceptual audio coder
    Rodrigues, JM
    Tomé, AM
    IEEE TRANSACTIONS ON SPEECH AND AUDIO PROCESSING, 2000, 8 (04): : 488 - 490
  • [3] A new subband perceptual audio coder using CELP
    van der Vrecken, O
    Hubaut, L
    Coulon, F
    PROCEEDINGS OF THE 1998 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING, VOLS 1-6, 1998, : 3661 - 3664
  • [4] Optimal perceptual binary allocation for AAC audio coder
    Perreau-Guimaraes, M
    Bonnet, M
    Moreau, N
    GLOBECOM'99: SEAMLESS INTERCONNECTION FOR UNIVERSAL SERVICES, VOL 1-5, 1999, : 2198 - 2202
  • [5] Toward a perceptual theory of transparency
    Singh, M
    Anderson, BL
    PSYCHOLOGICAL REVIEW, 2002, 109 (03) : 492 - 519
  • [6] PEAQ-based psychoacoustic model for perceptual audio coder
    Hu, XP
    He, GM
    Hou, XP
    8th International Conference on Advanced Communication Technology, Vols 1-3: TOWARD THE ERA OF UBIQUITOUS NETWORKS AND SOCIETIES, 2006, : U1819 - U1823
  • [7] Fixed bit rate perceptual wavelet packet audio coder
    Gunawan, TS
    Ambikairajah, E
    Epps, J
    2004 9TH IEEE SINGAPORE INTERNATIONAL CONFERENCE ON COMMUNICATION SYSTEMS (ICCS), 2004, : 235 - 239
  • [8] Analysis and application of perceptual weighting for AVS-M audio coder
    Yang Yuhong
    Hu Ruimin
    Zhang Yong
    Zhang Wei
    2007 INTERNATIONAL CONFERENCE ON WIRELESS COMMUNICATIONS, NETWORKING AND MOBILE COMPUTING, VOLS 1-15, 2007, : 2923 - 2926
  • [9] Toward a perceptual theory of transparency.
    Singh, M
    Anderson, BL
    INVESTIGATIVE OPHTHALMOLOGY & VISUAL SCIENCE, 2000, 41 (04) : S218 - S218
  • [10] Performance evaluation of an audio perceptual subband coder with dynamic bit allocation
    Caini, C
    Coralli, AV
    DSP 97: 1997 13TH INTERNATIONAL CONFERENCE ON DIGITAL SIGNAL PROCESSING PROCEEDINGS, VOLS 1 AND 2: SPECIAL SESSIONS, 1997, : 567 - 570