Margin-based Parallel Corpus Mining with Multilingual Sentence Embeddings

被引:0
|
作者
Artetxe, Mikel [1 ]
Schwenk, Holger [2 ]
机构
[1] Univ Basque Country, UPV EHU, Leioa, Spain
[2] Facebook AI Res, Menlo Pk, CA USA
关键词
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Machine translation is highly sensitive to the size and quality of the training data, which has led to an increasing interest in collecting and filtering large parallel corpora. In this paper, we propose a new method for this task based on multilingual sentence embeddings. In contrast to previous approaches, which rely on nearest neighbor retrieval with a hard threshold over cosine similarity, our proposed method accounts for the scale inconsistencies of this measure, considering the margin between a given sentence pair and its closest candidates instead. Our experiments show large improvements over existing methods. We outperform the best published results on the BUCC mining task and the UN reconstruction task by more than 10 F1 and 30 precision points, respectively. Filtering the English-German ParaCrawl corpus with our approach, we obtain 31.2 BLEU points on newstest2014, an improvement of more than one point over the best official filtered version.
引用
收藏
页码:3197 / 3203
页数:7
相关论文
共 50 条
  • [41] A note on margin-based loss functions in classification
    Lin, Y
    [J]. STATISTICS & PROBABILITY LETTERS, 2004, 68 (01) : 73 - 82
  • [42] Margin-based ranking meets boosting in the middle
    Rudin, C
    Cortes, C
    Mohri, M
    Schapire, RE
    [J]. LEARNING THEORY, PROCEEDINGS, 2005, 3559 : 63 - 78
  • [43] Margin-Based Discriminative Training for String Recognition
    Heigold, Georg
    Dreuw, Philippe
    Hahn, Stefan
    Schlueter, Ralf
    Ney, Hermann
    [J]. IEEE JOURNAL OF SELECTED TOPICS IN SIGNAL PROCESSING, 2010, 4 (06) : 917 - 925
  • [44] On Margin-Based Cluster Recovery with Oracle Queries
    Bressan, Marco
    Cesa-Bianchi, Nicolo
    Lattanzi, Silvio
    Paudice, Andrea
    [J]. ADVANCES IN NEURAL INFORMATION PROCESSING SYSTEMS 34 (NEURIPS 2021), 2021, 34
  • [45] Cross-Lingual Classification of Political Texts Using Multilingual Sentence Embeddings
    Licht, Hauke
    [J]. POLITICAL ANALYSIS, 2023, 31 (03): : 366 - 379
  • [46] Unified Binary and Multiclass Margin-Based Classification
    Wang, Yutong
    Scott, Clayton
    [J]. JOURNAL OF MACHINE LEARNING RESEARCH, 2024, 25 : 1 - 51
  • [47] ASYMPTOTIC BEHAVIOR OF MARGIN-BASED CLASSIFICATION METHODS
    Huang, Hanwen
    [J]. 2018 IEEE STATISTICAL SIGNAL PROCESSING WORKSHOP (SSP), 2018, : 463 - 467
  • [48] Margin-based active learning for structured predictions
    Kevin Small
    Dan Roth
    [J]. International Journal of Machine Learning and Cybernetics, 2010, 1 : 3 - 25
  • [49] Margin-based active learning for LVQ networks
    Schleif, F-M.
    Hammer, B.
    Villmann, T.
    [J]. NEUROCOMPUTING, 2007, 70 (7-9) : 1215 - 1224
  • [50] Efficient Large Margin-Based Feature Extraction
    Zhao, Guodong
    Wu, Yan
    [J]. NEURAL PROCESSING LETTERS, 2019, 50 (02) : 1257 - 1279