Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings

被引:0
|
作者
Artetxe, Mikel [1 ]
Schwenk, Holger [2 ]
机构
[1] Univ Basque Country, UPV EHU, Leioa, Spain
[2] Facebook AI Res, Menlo Pk, CA USA
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for this task based on multilingual sentence embeddings. In contrast to previous approaches, which rely on nearest neighbor retrieval with a hard threshold over cosine similarity, our proposed method accounts for the scale inconsistencies of this measure, considering the margin between a given sentence pair and its closest candidates instead. Our experiments show large improvements over existing methods. We outperform the best published results on the BUCC mining task and the UN reconstruction task by more than 10 F1 and 30 precision points, respectively. Filtering the English-German ParaCrawl corpus with our approach, we obtain 31.2 BLEU points on newstest2014, an improvement of more than one point over the best official filtered version.
引用
收藏
页码:3197 / 3203
页数:7
相关论文
共 50 条
  • [1] Unsupervised Multilingual Sentence Embeddings for Parallel Corpus Mining
    Kvapilikova, Ivana
    Artetxe, Mikel
    Labaka, Gorka
    Agirre, Eneko
    Bojar, Ondrej
    [J]. 58TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS (ACL 2020): STUDENT RESEARCH WORKSHOP, 2020, : 255 - 262
  • [2] Learning Multilingual Sentence Embeddings from Monolingual Corpus
    Wang, Shuai
    Hou, Lei
    Li, Juanzi
    Tong, Meihan
    Jiang, Jiabo
    [J]. CHINESE COMPUTATIONAL LINGUISTICS, CCL 2019, 2019, 11856 : 346 - 357
  • [3] Low-Resource Corpus Filtering using Multilingual Sentence Embeddings
    Chaudhary, Vishrav
    Tang, Yuqing
    Guzman, Francisco
    Schwenk, Holger
    Koehn, Philipp
    [J]. FOURTH CONFERENCE ON MACHINE TRANSLATION (WMT 2019), VOL 3: SHARED TASK PAPERS, DAY 2, 2019, : 261 - 266
  • [4] Are the Best Multilingual Document Embeddings simply Based on Sentence Embeddings?
    Sannigrahi, Sonal
    van Genabith, Josef
    Espana-Bonet, Cristina
    [J]. 17TH CONFERENCE OF THE EUROPEAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, EACL 2023, 2023, : 2306 - 2316
  • [5] Learning Deep Embeddings via Margin-Based Discriminate Loss
    Sun, Peng
    Tang, Wenzhong
    Bai, Xiao
    [J]. STRUCTURAL, SYNTACTIC, AND STATISTICAL PATTERN RECOGNITION, S+SSPR 2018, 2018, 11004 : 107 - 115
  • [6] Margin-based active learning and background knowledge in text mining
    Silva, C
    Ribeiro, B
    [J]. HIS'04: FOURTH INTERNATIONAL CONFERENCE ON HYBRID INTELLIGENT SYSTEMS, PROCEEDINGS, 2005, : 8 - 13
  • [7] Emu: Enhancing Multilingual Sentence Embeddings with Semantic Specialization
    Hirota, Wataru
    Suhara, Yoshihiko
    Golshan, Behzad
    Tan, Wang-Chiew
    [J]. THIRTY-FOURTH AAAI CONFERENCE ON ARTIFICIAL INTELLIGENCE, THE THIRTY-SECOND INNOVATIVE APPLICATIONS OF ARTIFICIAL INTELLIGENCE CONFERENCE AND THE TENTH AAAI SYMPOSIUM ON EDUCATIONAL ADVANCES IN ARTIFICIAL INTELLIGENCE, 2020, 34 : 7935 - 7943
  • [8] Multilingual sentence categorization and novelty mining
    Zhang, Yi
    Tsai, Flora S.
    Kwee, Agus Trisnajaya
    [J]. INFORMATION PROCESSING & MANAGEMENT, 2011, 47 (05) : 667 - 675
  • [9] Margin-Based Transfer Learning
    Su, Bai
    Xu, Wei
    Shen, Yidong
    [J]. SIXTH INTERNATIONAL SYMPOSIUM ON NEURAL NETWORKS (ISNN 2009), 2009, 56 : 223 - +
  • [10] COMFO: Multilingual Corpus for Opinion Mining
    Faty, Lamine
    Drame, Khadim
    Sarr, Edouard Ngor
    Ndiaye, Marie
    Diop, Ibrahima
    Dia, Yoro
    Sall, Ousmane
    [J]. ARTIFICIAL GENERAL INTELLIGENCE, AGI 2022, 2023, 13539 : 14 - 19