Representation of structured data of the text genre as a technique for automatic text processing

被引:0
|
作者
Fonseca, Claudia Aparecida [1 ]
Carvalho Guelpeli, Marcus Vinicius [2 ]
de Souza Netto, Rafael Santiago [3 ]
机构
[1] Univ Fed Vales Jequitinhonha & Mucuri, Dept Letras, Diamantina, MG, Brazil
[2] Univ Fed Vales Jequitinhonha & Mucuri, Dept Sistema Informacao, Diamantina, MG, Brazil
[3] Ctr Univ Barra Mansa, Dept Ciencia Comp, Barra Mansa, Rio De Janeiro, Brazil
来源
关键词
Corpus linguistics; Natural language processing; Scientific article; Text genre; Corpora annotation;
D O I
10.35699/1983-3652.2021.35445
中图分类号
H [语言、文字];
学科分类号
05 ;
摘要
The present article was developed in the field of Natural Language Processing and Language Studies based on a corpus compiled by computational tools. This study is based on the assumption that it is helpful to trace a close relationship between corpus generation/annotation and the assessment of the constitutive elements of the text genre source. It aims to demonstrate, through specific studies of structured data from the text genre 'scientific article', alternatives to automatic text processing techniques. In order to reach the intended goal, the authors created a computational model for the compilation of a linguistic, specialized Corpus, representative of the genre Scientific Article -CorpACE. The object of study includes the constitutive elements of scientific articles, marked in XML, extracted and collected from the SciELO-Scientific Electronic Library On-line database. The final product was a database obtained with information extracted and structured in XML format, which designates and identifies the markups of the genre being analyzed and is available for many tools and applications. The results demonstrate how the representation of constitutive elements of the genre can condense available information with hierarchical and dynamic processes built during the compilation. At the end of the study, it is believed that more research will be required for bringing Language Science and Computer Science closer with emphasis on NLP in the attempt to represent and manipulate linguistic knowledge in its many levels - morphological, syntactic, semantic and discursive - in order to improve implementation and manipulation of automatic text processing.
引用
收藏
页数:26
相关论文
共 50 条
  • [1] Automatic detection of text genre
    Kessler, B
    Nunberg, G
    Schutze, H
    [J]. 35TH ANNUAL MEETING OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS AND THE 8TH CONFERENCE OF THE EUROPEAN CHAPTER OF THE ASSOCIATION FOR COMPUTATIONAL LINGUISTICS, PROCEEDINGS OF THE CONFERENCE, 1997, : 32 - 38
  • [2] Proposed Architecture for Automatic Conversion of Unstructured Text Data into Structured Text Data on the Web
    Madhusudhan, Ch.
    Rao, K. Mrithyunjaya
    [J]. INTERNATIONAL JOURNAL OF COMPUTER SCIENCE AND NETWORK SECURITY, 2013, 13 (12): : 110 - 116
  • [3] Lightweight structured text processing
    Miller, RC
    Myers, BA
    [J]. PROCEEDINGS OF THE 1999 USENIX ANNUAL TECHNICAL CONFERENCE, 1999, : 131 - 144
  • [4] Automatic Genre Recognition and Adaptive Text Summarization
    Yatsko, V. A.
    Starikov, M. S.
    Butakov, A. V.
    [J]. AUTOMATIC DOCUMENTATION AND MATHEMATICAL LINGUISTICS, 2010, 44 (03) : 111 - 120
  • [5] Automatic text categorization in terms of genre and author
    Stamatatos, E
    Kokkinakis, G
    Fakotakis, N
    [J]. COMPUTATIONAL LINGUISTICS, 2000, 26 (04) : 471 - 495
  • [6] Text Incoherence, or Some Pitfalls of Automatic Text Processing
    Inkova, O. Yu
    [J]. VESTNIK TOMSKOGO GOSUDARSTVENNOGO UNIVERSITETA FILOLOGIYA-TOMSK STATE UNIVERSITY JOURNAL OF PHILOLOGY, 2021, 74 : 81 - 98
  • [7] Automatic segmentation of text into structured records
    Borkar, V
    Deshmukh, K
    Sarawagi, S
    [J]. SIGMOD RECORD, 2001, 30 (02) : 175 - 186
  • [8] Automatic Processing of Arabic Text
    Osman, Ziad
    Hamandi, Lama
    Zantout, Rached
    Sibai, Fadi N.
    [J]. 2009 INTERNATIONAL CONFERENCE ON INNOVATIONS IN INFORMATION TECHNOLOGY, 2009, : 6 - +
  • [9] Graphical Models for Text: A New Paradigm for Text Representation and Processing
    Aggarwal, Charu C.
    Zhao, Peixiang
    [J]. SIGIR 2010: PROCEEDINGS OF THE 33RD ANNUAL INTERNATIONAL ACM SIGIR CONFERENCE ON RESEARCH DEVELOPMENT IN INFORMATION RETRIEVAL, 2010, : 899 - 900
  • [10] GENRE AS TEXT
    PEREZFIRMAT, G
    [J]. COMPARATIVE LITERATURE STUDIES, 1980, 17 (01) : 16 - 25