Topic-Based Coherence Modeling for Statistical Machine Translation

被引:20
|
作者
Xiong, Deyi [1 ,2 ]
Zhang, Min [1 ,2 ]
Wang, Xing [2 ]
机构
[1] Inst Infocomm Res, Singapore 138632, Singapore
[2] Soochow Univ, Prov Key Lab Comp Informat Proc Technol, Suzhou 215006, Peoples R China
基金
中国国家自然科学基金;
关键词
Text coherence; text analysis; coherence chain; topic modeling; statistical machine translation (SMT); natural language processing;
D O I
10.1109/TASLP.2015.2395254
中图分类号
O42 [声学];
学科分类号
070206 ; 082403 ;
摘要
Coherence that ties sentences of a text into a meaningfully connected structure is of great importance to text generation and translation. In this paper, we propose topic-based coherence models to produce coherence for document translation, in terms of the continuity of sentence topics in a text. We automatically extract a coherence chain for each source text to be translated. Based on the extracted source coherence chain, we adopt a maximum entropy classifier to predict the target coherence chain that defines a linear topic structure for the target document. We build two topic-based coherence models on the predicted target coherence chain: 1) a word level coherence model that helps the decoder select coherent word translations and 2) a phrase level coherence model that guides the decoder to select coherent phrase translations. We integrate the two models into a state-of-the-art phrase-based machine translation system. Experiments on large-scale training data show that our coherence models achieve substantial improvements over both the baseline and models that are built on either document topics or sentence topics obtained under the assumption of direct topic correspondence between the source and target side. Additionally, further evaluations on translation outputs suggest that target translations generated by our coherence models are more coherent and similar to reference translations than those generated by the baseline.
引用
收藏
页码:483 / 493
页数:11
相关论文
共 50 条
  • [1] Topic-based coherence modeling for statistical machine translation
    Institute for Infocomm Research, Singapore
    138632, Singapore
    不详
    215006, China
    [J]. IEEE Trans. Audio Speech Lang. Process., 3 (483-493):
  • [2] Topic-based term translation models for statistical machine translation
    Xiong, Deyi
    Meng, Fandong
    Liu, Qun
    [J]. ARTIFICIAL INTELLIGENCE, 2016, 232 : 54 - 75
  • [3] Topic Adaptation for Statistical Machine Translation
    Taraghi, Mina
    Khadivi, Shahram
    [J]. 2017 25TH IRANIAN CONFERENCE ON ELECTRICAL ENGINEERING (ICEE), 2017, : 2147 - 2152
  • [4] A Topic-Triggered Translation Model for Statistical Machine Translation
    SU Jinsong
    WANG Zhihao
    WU Qingqiang
    YAO Junfeng
    LONG Fei
    ZHANG Haiying
    [J]. Chinese Journal of Electronics, 2017, 26 (01) : 65 - 72
  • [5] A Topic-Triggered Translation Model for Statistical Machine Translation
    Su Jinsong
    Wang Zhihao
    Wu Qingqiang
    Yao Junfeng
    Long Fei
    Zhang Haiying
    [J]. CHINESE JOURNAL OF ELECTRONICS, 2017, 26 (01) : 65 - 72
  • [6] Topic-Based Dissimilarity and Sensitivity Models for Translation Rule Selection
    Zhang, Min
    Xiao, Xinyan
    Xiong, Deyi
    Liu, Qun
    [J]. JOURNAL OF ARTIFICIAL INTELLIGENCE RESEARCH, 2014, 50 : 1 - 30
  • [7] Evaluation of topic-based adaptation and student modeling in QuizGuide
    Sergey Sosnovsky
    Peter Brusilovsky
    [J]. User Modeling and User-Adapted Interaction, 2015, 25 : 371 - 424
  • [8] Evaluation of topic-based adaptation and student modeling in QuizGuide
    Sosnovsky, Sergey
    Brusilovsky, Peter
    [J]. USER MODELING AND USER-ADAPTED INTERACTION, 2015, 25 (04) : 371 - 424
  • [9] Topic-Based Language Modeling with Dynamic Bayesian Networks
    Wiggers, Pascal
    Rothkrantz, Leon J. M.
    [J]. INTERSPEECH 2006 AND 9TH INTERNATIONAL CONFERENCE ON SPOKEN LANGUAGE PROCESSING, VOLS 1-5, 2006, : 1866 - 1869
  • [10] Improving statistical machine translation using topic information
    Yu, Hui
    Wang, Xin
    Shao, Zengzhen
    Xu, Weizhi
    [J]. IPPTA: Quarterly Journal of Indian Pulp and Paper Technical Association, 2018, 30 (01) : 148 - 153