Linguistically Driven Multi-Task Pre-Training for Low-Resource Neural Machine Translation

被引：6

作者：

Mao, Zhuoyuan ^{[1
]}

Chu, Chenhui ^{[1
]}

Kurohashi, Sadao ^{[1
]}

机构：

[1] Kyoto Univ, Grad Sch Informat, Kyoto, Japan

来源：

ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING | 2022年 / 21卷 / 04期

关键词：

Low-resource neural machine translation; pre-training; linguistically-driven;

D O I：

10.1145/3491065

中图分类号：

TP18 [人工智能理论];

学科分类号：

081104 ; 0812 ; 0835 ; 1405 ;

摘要：

In the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese-English & Japanese-Chinese, Wikipedia Japanese-Chinese, News English-Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese-English tasks, up to +7.0 BLEU points for the Japanese-Chinese tasks and up to +1.3 BLEU points for English-Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency.

引用

页数：29

共 50 条

[41] Machine Translation in Low-Resource Languages by an Adversarial Neural Network
Sun, Mengtao
Wang, Hao
Pasquine, Mark
Hameed, Ibrahim A.
APPLIED SCIENCES-BASEL, 2021, 11 (22):
[42] A Strategy for Referential Problem in Low-Resource Neural Machine Translation
Ji, Yatu
Shi, Lei
Su, Yila
Ren, Qing-dao-er-ji
Wu, Nier
Wang, Hongbin
ARTIFICIAL NEURAL NETWORKS AND MACHINE LEARNING, ICANN 2021, PT V, 2021, 12895 : 321 - 332
[43] A Multi-lingual Multi-task Architecture for Low-resource Sequence Labeling
Lin, Ying
Yang, Shengqi
Stoyanov, Veselin
Ji, Heng
PROCEEDINGS OF THE 56TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL), VOL 1, 2018, : 799 - 809
[44] Unsupervised Source Hierarchies for Low-Resource Neural Machine Translation
Currey, Anna
Heafield, Kenneth
RELEVANCE OF LINGUISTIC STRUCTURE IN NEURAL ARCHITECTURES FOR NLP, 2018, : 6 - 12
[45] Language Model Prior for Low-Resource Neural Machine Translation
Baziotis, Christos
Haddow, Barry
Birch, Alexandra
PROCEEDINGS OF THE 2020 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP), 2020, : 7622 - 7634
[46] Low-Resource Neural Machine Translation: A Systematic Literature Review
Yazar, Bilge Kagan
Sahin, Durmus Ozkan
Kilic, Erdal
IEEE ACCESS, 2023, 11 : 131775 - 131813
[47] Meta-Learning for Low-Resource Neural Machine Translation
Gu, Jiatao
Wang, Yong
Chen, Yun
Cho, Kyunghyun
Li, Victor O. K.
2018 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP 2018), 2018, : 3622 - 3631
[48] Universal Conditional Masked Language Pre-training for Neural Machine Translation
Li, Pengfei
Li, Liangyou
Zhang, Meng
Wu, Minghao
Liu, Qun
PROCEEDINGS OF THE 60TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2022), VOL 1: (LONG PAPERS), 2022, : 6379 - 6391
[49] Pre-training Multilingual Neural Machine Translation by Leveraging Alignment Information
Lin, Zehui
Pan, Xiao
Wang, Mingxuan
Qiu, Xipeng
Feng, Jiangtao
Zhou, Hao
Li, Lei
PROCEEDINGS OF THE 2020 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING (EMNLP), 2020, : 2649 - 2663
[50] Revisiting Low-Resource Neural Machine Translation: A Case Study
Sennrich, Rico
Zhang, Biao
57TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2019), 2019, : 211 - 221

← 1 2 3 4 5 →