The left and right context of a word: Overlapping chinese syllable word segmentation with minimal context

被引:0
|
作者
Jiang, Mike Tian-Jian [1 ,4 ]
Lee, Tsung-Hsien [2 ,5 ]
Hsu, Wen-Lian [3 ,5 ]
机构
[1] National Tsing Hua University and Academia Sinica, Taiwan
[2] Academia Sinica and University of Texas, Austin, United States
[3] Academia Sinica and National Tsing Hua University, Taiwan
[4] Computer Science Department, National Tsing Hua University, Taiwan
[5] Institute of Information Science, Academia Sinica, Taiwan
关键词
Computing power - Linguistics - Random processes - Natural language processing systems;
D O I
10.1145/2425327.2425329
中图分类号
学科分类号
摘要
Since a Chinese syllable can correspond to many characters (homophones), the syllable-to-character conversion task is quite challenging for Chinese phonetic input methods (CPIM). There are usually two stages in a CPIM: 1. segment the syllable sequence into syllable words, and 2. select the most likely character words for each syllable word. A CPIM usually assumes that the input is a complete sentence, and evaluates the performance based on a well-formed corpus. However, in practice, most Pinyin users prefer progressive text entry in several short chunks, mainly in one or two words each (most Chinese words consist of two or more characters). Short chunks do not provide enough contexts to perform the best possible syllable-to-character conversion, especially when a chunk consists of overlapping syllable words. In such cases, a conversion system often selects the boundary of a word with the highest frequency. Short chunk input is even more popular on platforms with limited computing power, such as mobile phones. Based on the observation that the relative strength of a word can be quite different when calculated leftwards or rightwards, we propose a simple division of the word context into the left context and the right context. Furthermore, we design a double ranking strategy for each word to reduce the number of errors in Step 1. Our strategy is modeled as the minimum feedback arc set problem on bipartite tournament with approximate solutions derived from genetic algorithm. Experiments show that, compared to the frequency-based method (FBM) (low memory and fast) and the conditional random fields (CRF) model (larger memory and slower), our double ranking strategy has the benefits of less memory and low power requirement with competitive performance. We believe a similar strategy could also be adopted to disambiguate conflicting linguistic patterns effectively. Copyright © 2013 ACM.
引用
收藏
相关论文
共 50 条
  • [1] Chinese word segmentation with local and global context representation learning
    School of Computer and Communication Engineering, University of Science and Technology Beijing, Beijing
    100083, China
    不详
    100190, China
    High Technol Letters, 1 (71-77):
  • [2] The Research of Chinese Word Segmentation Disambiguation Dased on Context Information
    Mai Fanjin
    Le, Zhao
    Ling, Huang
    ICIIP'18: PROCEEDINGS OF THE 3RD INTERNATIONAL CONFERENCE ON INTELLIGENT INFORMATION PROCESSING, 2018, : 101 - 106
  • [3] Chinese word segmentation with local and global context representation learning
    李岩
    Zhang Yinghua
    Huang Xiaoping
    Yin Xucheng
    Hao Hongwei
    High Technology Letters, 2015, 21 (01) : 71 - 77
  • [4] Analysis on Effect Range of Context in Chinese Word Segmentation based Word-position Tagging
    Wang, Xijie
    Guo, An
    2012 FOURTH INTERNATIONAL CONFERENCE ON MULTIMEDIA INFORMATION NETWORKING AND SECURITY (MINES 2012), 2012, : 552 - 555
  • [5] Context-based Approach for Covering Ambiguity Resolution in Chinese Word Segmentation
    Feng, Su-qin
    Hou, Su-qin
    ICIC 2009: SECOND INTERNATIONAL CONFERENCE ON INFORMATION AND COMPUTING SCIENCE, VOL 2, PROCEEDINGS: IMAGE ANALYSIS, INFORMATION AND SIGNAL PROCESSING, 2009, : 43 - +
  • [6] The Translatology Process of the Word ‘Gaze’ in Chinese Context
    Tong JunForeign Language DepartmentGuizhou UniversityGuiyangGuizhou
    科技信息(科学教研), 2008, (01) : 260 - 260
  • [7] Word Segmentation of Overlapping Ambiguous Strings During Chinese Reading
    Ma, Guojie
    Li, Xingshan
    Rayner, Keith
    JOURNAL OF EXPERIMENTAL PSYCHOLOGY-HUMAN PERCEPTION AND PERFORMANCE, 2014, 40 (03) : 1046 - 1059
  • [8] Zipfian frequency distributions facilitate word segmentation in context
    Kurumada, Chigusa
    Meylan, Stephan C.
    Frank, Michael C.
    COGNITION, 2013, 127 (03) : 439 - 453
  • [9] A Bayesian framework for word segmentation: Exploring the effects of context
    Goldwater, Sharon
    Griffiths, Thomas L.
    Johnson, Mark
    COGNITION, 2009, 112 (01) : 21 - 54
  • [10] Chinese compound word inference through context and word-internal cues
    Xu, Yi
    Zhang, Jie
    LANGUAGE TEACHING RESEARCH, 2022, 26 (03) : 308 - 332