Automatic Category Theme Identification and Hierarchy Generation for Chinese Text Categorization

被引:0
|
作者
Hsin-Chang Yang
Chung-Hong Lee
机构
[1] Chang Jung University,Department of Information Management
[2] National Kaohsiung University of Applied Sciences,Department of Electrical Engineering
关键词
automatic category theme identification; automatic category hierarchy generation; text categorization; self-organizing maps; text mining;
D O I
暂无
中图分类号
学科分类号
摘要
Recently research on text mining has attracted lots of attention from both industrial and academic fields. Text mining concerns of discovering unknown patterns or knowledge from a large text repository. The problem is not easy to tackle due to the semi-structured or even unstructured nature of those texts under consideration. Many approaches have been devised for mining various kinds of knowledge from texts. One important aspect of text mining is on automatic text categorization, which assigns a text document to some predefined category if the document falls into the theme of the category. Traditionally the categories are arranged in hierarchical manner to achieve effective searching and indexing as well as easy comprehension for human beings. The determination of category themes and their hierarchical structures were most done by human experts. In this work, we developed an approach to automatically generate category themes and reveal the hierarchical structure among them. We also used the generated structure to categorize text documents. The document collection was trained by a self-organizing map to form two feature maps. These maps were then analyzed to obtain the category themes and their structure. Although the test corpus contains documents written in Chinese, the proposed approach can be applied to documents written in any language and such documents can be transformed into a list of separated terms.
引用
收藏
页码:47 / 67
页数:20
相关论文
共 50 条
  • [1] Automatic category theme identification and hierarchy generation for Chinese text categorization
    Yang, HC
    Lee, CH
    [J]. JOURNAL OF INTELLIGENT INFORMATION SYSTEMS, 2005, 25 (01) : 47 - 67
  • [2] Automatic Category Structure Generation and Categorization of Chinese Text Documents
    Yang, Hsin-Chang
    Lee, Chung-Hong
    [J]. LECTURE NOTES IN COMPUTER SCIENCE <D>, 2000, 1910 : 673 - 678
  • [3] The Chinese text categorization system with association rule and category priority
    Chiang, Ding-An
    Keh, Huan-Chao
    Huang, Hui-Hua
    Chyr, Derming
    [J]. EXPERT SYSTEMS WITH APPLICATIONS, 2008, 35 (1-2) : 102 - 110
  • [4] Research on Chinese Text Automatic Categorization Based on VSM
    Tong Xiao-Jun
    Cui Ming-Gen
    Song Guo-Long
    [J]. 2007 INTERNATIONAL CONFERENCE ON WIRELESS COMMUNICATIONS, NETWORKING AND MOBILE COMPUTING, VOLS 1-15, 2007, : 3863 - +
  • [5] Exploiting hierarchy in text categorization
    Weigend A.S.
    Wiener E.D.
    Pedersen J.O.
    [J]. Information Retrieval, 1999, 1 (3): : 193 - 216
  • [6] Automatic Chinese Text Categorization System Based on Mutual Information
    Lu, Zhimao
    Shi, Hong
    Zhang, Qi
    Yuan, Chaoyue
    [J]. 2009 IEEE INTERNATIONAL CONFERENCE ON MECHATRONICS AND AUTOMATION, VOLS 1-7, CONFERENCE PROCEEDINGS, 2009, : 4986 - 4990
  • [7] Automatic Maintenance of the Category Hierarchy
    He, Lei
    Sun, Xiaoping
    [J]. 2013 NINTH INTERNATIONAL CONFERENCE ON SEMANTICS, KNOWLEDGE AND GRIDS (SKG), 2013, : 218 - 221
  • [8] Automatic maintenance of category hierarchy
    Hai Zhuge
    Lei He
    [J]. FUTURE GENERATION COMPUTER SYSTEMS-THE INTERNATIONAL JOURNAL OF ESCIENCE, 2017, 67 : 1 - 12
  • [9] Automatic category generation for text documents by self-organizing maps
    Yang, HC
    Lee, CH
    [J]. IJCNN 2000: PROCEEDINGS OF THE IEEE-INNS-ENNS INTERNATIONAL JOINT CONFERENCE ON NEURAL NETWORKS, VOL III, 2000, : 581 - 586
  • [10] Automatic generation of text categorization rules in a hybrid method based on machine learning
    Lana-Serrano, Sara
    Villena-Roman, Julio
    Collada-Perez, Sonia
    Carlos Gonzalez-Cristobal, Jose
    [J]. PROCESAMIENTO DEL LENGUAJE NATURAL, 2011, (47): : 231 - 237