A Data-Driven Model for Automated Chinese Word Segmentation and POS Tagging

被引:0
|
作者
Xu, Qing [1 ]
Wang, Zhiyou [2 ]
机构
[1] Changsha Univ Sci & Technol, Changsha 410000, Hunan, Peoples R China
[2] Changsha Univ, Sch Elect Commun & Elect Engn, Changsha 410000, Hunan, Peoples R China
关键词
NEURAL-NETWORK; EXTRACTION;
D O I
10.1155/2022/7622392
中图分类号
Q [生物科学];
学科分类号
07 ; 0710 ; 09 ;
摘要
Chinese natural language processing tasks often require the solution of Chinese word segmentation and POS tagging problems. Traditional Chinese word segmentation and POS tagging methods mainly use simple matching algorithms based on lexicons and rules. The simple matching or statistical analysis requires manual word segmentation followed by POS tagging, which leads to the inability to meet the practical requirements for label prediction accuracy. With the continuous development of deep learning technology, data-driven machine learning models provide new opportunities for automated Chinese word segmentation and POS tagging. Therefore, a data-driven automated Chinese word segmentation and POS tagging model is proposed in order to address the above problems. Firstly, the main idea and overall framework of the proposed automated model are outlined, and the tagging strategy and neural network language model used are described. Secondly, two main optimisations are made on the input side of the model: (1) the use of word2Vec for the representation of text features, thus representing the text as a distributed word vector; and (2) the use of an improved AlexNet for efficient encoding of long-range word, and the addition of an attention mechanism to the model. Finally, on the output side, an additional auxiliary loss function was designed to optimise the Chinese text based on its frequency. The experimental results show that the proposed model can significantly improve the accuracy and operational efficiency of Chinese word segmentation and POS tagging compared with other existing models, thus verifying its effectiveness and advancement.
引用
收藏
页数:10
相关论文
共 50 条
  • [1] An Effective Joint Model for Chinese Word Segmentation and POS Tagging
    Wang, Heng-Jun
    Si, Nian-Wen
    Chen, Cheng
    [J]. PROCEEDINGS OF THE 2016 INTERNATIONAL CONFERENCE ON INTELLIGENT INFORMATION PROCESSING (ICIIP'16), 2016,
  • [2] Word segmentation and POS tagging for Chinese keyphrase extraction
    Huang, XC
    Chen, J
    Yan, PL
    Luo, X
    [J]. ADVANCED DATA MINING AND APPLICATIONS, PROCEEDINGS, 2005, 3584 : 364 - 369
  • [3] Joint Chinese Word Segmentation and POS Tagging Using an Error-Driven Word-Character Hybrid Model
    Kruengkrai, Canasai
    Uchimoto, Kiyotaka
    Kazama, Jun'ichi
    Wang, Yiou
    Torisawa, Kentaro
    Isahara, Hitoshi
    [J]. IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2009, E92D (12) : 2298 - 2305
  • [4] A Unified Model for Joint Chinese Word Segmentation and POS Tagging with Heterogeneous Annotation Corpora
    Zhao, Jiayi
    Qiu, Xipeng
    Huang, Xuanjing
    [J]. 2013 INTERNATIONAL CONFERENCE ON ASIAN LANGUAGE PROCESSING (IALP 2013), 2013, : 227 - 230
  • [5] Comparing data-driven learning algorithms for PoS tagging of Swedish
    Megyesi, B
    [J]. PROCEEDINGS OF THE 2001 CONFERENCE ON EMPIRICAL METHODS IN NATURAL LANGUAGE PROCESSING, 2001, : 151 - 158
  • [6] A hybrid approach to word segmentation and POS tagging
    Oki Electric Industry Co., Ltd., 2−5−7 Honmachi, Chuo-ku, Osaka
    541−0053, Japan
    不详
    619−0289, Japan
    [J]. Proc. Annu. Meet. Assoc. Comput Linguist., 1600, (217-220):
  • [7] Simple semi-supervised learning for chinese word segmentation and pos tagging
    Li, Xinxin
    Wang, Xuan
    Waqas, Muhammad
    Harbin, Anwar
    [J]. Information Technology Journal, 2013, 12 (20) : 5955 - 5961
  • [8] A Simple and Effective Neural Model for Joint Word Segmentation and POS Tagging
    Zhang, Meishan
    Yu, Nan
    Fu, Guohong
    [J]. IEEE-ACM TRANSACTIONS ON AUDIO SPEECH AND LANGUAGE PROCESSING, 2018, 26 (09) : 1528 - 1538
  • [9] Character-Level Dependency Model for Joint Word Segmentation, POS Tagging, and Dependency Parsing in Chinese
    Guo, Zhen
    Zhang, Yujie
    Su, Chen
    Xu, Jinan
    Isahara, Hitoshi
    [J]. IEICE TRANSACTIONS ON INFORMATION AND SYSTEMS, 2016, E99D (01): : 257 - 264
  • [10] A Neural Joint Model with BERT for Burmese Syllable Segmentation, Word Segmentation, and POS Tagging
    Mao, Cunli
    Man, Zhibo
    Yu, Zhengtao
    Gao, Shengxiang
    Wang, Zhenhan
    Wang, Hongbin
    [J]. ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING, 2021, 20 (04)