Attentive Multi-Layer Perceptron for Non-autoregressive Generation

被引:0
|
作者
Jiang, Shuyang [1 ]
Zhang, Jun [2 ]
Feng, Jiangtao [2 ]
Zheng, Lin [3 ]
Kong, Lingpeng [3 ]
机构
[1] Shanghai Jiao Tong Univ, Shanghai, Peoples R China
[2] Shanghai Artificial Intelligence Lab, Shanghai, Peoples R China
[3] Univ Hong Kong, Hong Kong, Peoples R China
关键词
AMLP; Multi-Layer Perceptron; Attention Mechanism; Non-Autoregressive Model; TRANSLATION;
D O I
10.1007/978-3-031-43415-0_36
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Autoregressive (AR) generation almost dominates sequence generation for its efficacy. Recently, non-autoregressive (NAR) generation gains increasing popularity for its efficiency and growing efficacy. However, its efficiency is still bottlenecked by quadratic complexity in sequence lengths, which is prohibitive for scaling to long sequence generation and few works have been done to mitigate this problem. In this paper, we propose a novel MLP variant, Attentive Multi-Layer Perceptron (AMLP), to produce a generation model with linear time and space complexity. Different from classic MLP with static and learnable projection matrices, AMLP leverages adaptive projections computed from inputs in an attentive mode. The sample-aware adaptive projections enable communications among tokens in a sequence, and model the measurement between the query and key space. Furthermore, we marry AMLP with popular NAR models, deriving a highly efficient NAR-AMLP architecture with linear time and space complexity. Empirical results show that such marriage architecture surpasses competitive efficient NAR models, by a significant margin on text-to-speech synthesis and machine translation. We also test AMLP's self- and cross-attention ability separately with extensive ablation experiments, and find them comparable or even superior to the other efficient models. The efficiency analysis further shows that AMLP extremely reduces the memory cost against vanilla non-autoregressive models for long sequences.
引用
收藏
页码:612 / 629
页数:18
相关论文
共 50 条
  • [41] Fuzzy multi-layer perceptron for binary pattern recognition
    Canuto, AMP
    Howells, WGJ
    Fairhurst, MC
    SEVENTH INTERNATIONAL CONFERENCE ON IMAGE PROCESSING AND ITS APPLICATIONS, 1999, (465): : 260 - 264
  • [42] Multi-layer perceptron based modelling of nonlinear systems
    Lightbody, G
    Irwin, GW
    FUZZY SETS AND SYSTEMS, 1996, 79 (01) : 93 - 112
  • [43] Prediction of zenith tropospheric delay by multi-layer perceptron
    Katsougiannopoulos, S.
    Pikridas, C.
    JOURNAL OF APPLIED GEODESY, 2009, 3 (04) : 223 - 229
  • [44] A novel scheme to determine the architecture of a Multi-layer Perceptron
    Chintalapudi, KK
    Pal, NR
    1998 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN, AND CYBERNETICS, VOLS 1-5, 1998, : 2297 - 2302
  • [45] Online phoneme recognition using multi-layer perceptron networks combined with recurrent non-linear autoregressive neural networks with exogenous inputs
    Bonilla Cardona, Diana A.
    Nedjah, Nadia
    Mourelle, Luiza M.
    NEUROCOMPUTING, 2017, 265 : 78 - 90
  • [46] Spot the Error: Non-autoregressive Graphic Layout Generation with Wireframe Locator
    Lin, Jieru
    Huang, Danqing
    Zhao, Tiejun
    Zhan, Dechen
    Lin, Chin-Yew
    THIRTY-EIGHTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, VOL 38 NO 4, 2024, : 3413 - 3421
  • [47] Non-autoregressive Translation with Layer-Wise Prediction and Deep Supervision
    Huang, Chenyang
    Zhou, Hao
    Zaiane, Osmar R.
    Mou, Lili
    Li, Lei
    THIRTY-SIXTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE / THIRTY-FOURTH CONFERENCE ON INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE / TWELVETH SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2022, : 10776 - 10784
  • [48] Sequence Labeling as Non-Autoregressive Dual-Query Set Generation
    Chen, Xiang
    Li, Lei
    Zhu, Yuqi
    Deng, Shumin
    Tan, Chuanqi
    Huang, Fei
    Si, Luo
    Zhang, Ningyu
    Chen, Huajun
    IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2024, 32 : 1546 - 1558
  • [49] Towards Personalized Bundle Creative Generation with Contrastive Non-Autoregressive Decoding
    Wei, Penghui
    Liu, Shaoguo
    Yang, Xuanhua
    Wang, Liang
    Zheng, Bo
    PROCEEDINGS OF THE 45TH INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH AND DEVELOPMENT IN INFORMATION RETRIEVAL (SIGIR '22), 2022, : 2634 - 2638
  • [50] A COMPARATIVE STUDY ON NON-AUTOREGRESSIVE MODELINGS FOR SPEECH-TO-TEXT GENERATION
    Higuchi, Yosuke
    Chen, Nanxin
    Fujita, Yuya
    Inaguma, Hirofumi
    Komatsu, Tatsuya
    Lee, Jaesong
    Nozaki, Jumon
    Wang, Tianzi
    Watanabe, Shinji
    2021 IEEE AUTOMATIC SPEECH RECOGNITION AND UNDERSTANDING WORKSHOP (ASRU), 2021, : 47 - 54