Part of speech tagging for Arabic

被引:7
|
作者
Kuebler, Sandra [1 ]
Mohamed, Emad [2 ]
机构
[1] Indiana Univ, Dept Linguist, Bloomington, IN 47405 USA
[2] Carnegie Mellon Univ Qatar Educ City, Comp Sci Program, Doha, Qatar
关键词
D O I
10.1017/S1351324911000325
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
This paper presents an investigation of part of speech (POS) tagging for Arabic as it occurs naturally, i.e. unvocalized text (without diacritics). We also do not assume any prior tokenization, although this was used previously as a basis for POS tagging. Arabic is a morphologically complex language, i.e. there is a high number of inflections per word; and the tagset is larger than the typical tagset for English. Both factors, the second one being partly dependent on the first, increase the number of word/tag combinations, for which the POS tagger needs to find estimates, and thus they contribute to data sparseness. We present a novel approach to Arabic POS tagging that does not require any pre-processing, such as segmentation or tokenization: whole word tagging. In this approach, the complete word is assigned a complex POS tag, which includes morphological information. A competing approach investigates the effect of segmentation and vocalization on POS tagging to alleviate data sparseness and ambiguity. In the segmentation-based approach, we first automatically segment words and then POS tags the segments. The complex tagset encompasses 993 POS tags, whereas the segment-based tagset encompasses only 139 tags. However, segments are also more ambiguous, thus there are more possible combinations of segment tags. In realistic situations, in which we have no information about segmentation or vocalization, whole word tagging reaches the highest accuracy of 94.74%. If gold standard segmentation or vocalization is available, including this information improves POS tagging accuracy. However, while our automatic segmentation and vocalization modules reach state-of-the-art performance, their performance is not reliable enough for POS tagging and actually impairs POS tagging performance. Finally, we investigate whether a reduction of the complex tagset to the Extra-Reduced Tagset as suggested by Habash and Rambow (Habash, N., and Rambow, O. 2005. Arabic tokenization, part-of-speech tagging and morphological disambiguation in one fell swoop. In Proceedings of the 43rd Annual Meeting of the Association for Computational Linguistics (ACL), Ann Arbor, MI, USA, pp. 573-80) will alleviate the data sparseness problem. While the POS tagging accuracy increases due to the smaller tagset, a closer look shows that using a complex tagset for POS tagging and then converting the resulting annotation to the smaller tagset results in a higher accuracy than tagging using the smaller tagset directly.
引用
收藏
页码:521 / 548
页数:28
相关论文
共 50 条
  • [1] Arabic Part of Speech Tagging
    Mohamed, Emad
    Kuebler, Sandra
    [J]. LREC 2010 - SEVENTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION, 2010, : 2537 - 2543
  • [2] Toward enhanced Arabic speech recognition using part of speech tagging
    AbuZeina, Dia
    Al-Khatib, Wasfi
    Elshafei, Moustafa
    Al-Muhtaseb, Husni
    [J]. INTERNATIONAL JOURNAL OF SPEECH TECHNOLOGY, 2011, 14 (04) : 419 - 426
  • [3] Morphological Segmentation and Part-of-Speech Tagging for the Arabic Heritage
    Mohamed, Emad
    [J]. ACM TRANSACTIONS ON ASIAN AND LOW-RESOURCE LANGUAGE INFORMATION PROCESSING, 2018, 17 (03)
  • [4] Parallel HMM-Based Approach for Arabic Part of Speech Tagging
    Kadim, Ayoub
    Lazrek, Azzeddine
    [J]. INTERNATIONAL ARAB JOURNAL OF INFORMATION TECHNOLOGY, 2018, 15 (02) : 341 - 351
  • [5] Improving Arabic Part-of-Speech Tagging through Morphological Analysis
    Albared, Mohammed
    Omar, Nazlia
    Ab Aziz, Mohd. Juzaiddin
    [J]. INTELLIGENT INFORMATION AND DATABASE SYSTEMS, ACIIDS 2011, PT I, 2011, 6591 : 317 - 326
  • [6] Part of speech tagging for Arabic text based radial basis function
    Shahin, Osama R.
    El Rwelli, Rady
    [J]. JOURNAL OF DISCRETE MATHEMATICAL SCIENCES & CRYPTOGRAPHY, 2021, 24 (08): : 2443 - 2459
  • [7] Part of Speech Tagging Approach to Designing Compound Words for Arabic Continuous Speech Recognition Systems
    AbuZeina, Dia
    Elshafei, Moustafa
    Al-Khatib, Wasfi
    [J]. INFORMATICS ENGINEERING AND INFORMATION SCIENCE, PT IV, 2011, 254 : 330 - 338
  • [8] Pattern-based Algorithm for Part-of-Speech Tagging Arabic Text
    Alqrainy, Shihadeh
    Alserhan, Hasan Muaidi
    Ayesh, Aladdin
    [J]. ICCES: 2008 INTERNATIONAL CONFERENCE ON COMPUTER ENGINEERING & SYSTEMS, 2007, : 119 - +
  • [9] Rule Based Approach for Arabic Part of Speech Tagging and Name Entity Recognition
    Btoush, Mohammad Hjouj
    Alarabeyyat, Abdulsalam
    Olab, Isa
    [J]. INTERNATIONAL JOURNAL OF ADVANCED COMPUTER SCIENCE AND APPLICATIONS, 2016, 7 (06) : 331 - 335
  • [10] Part-of-Speech Tagging for Arabic Gulf Dialect Using Bi-LSTM
    Alharbi, Randah
    Magdy, Walid
    Darwish, Kareem
    AbdelAli, Ahmed
    Mubarak, Hamdy
    [J]. PROCEEDINGS OF THE ELEVENTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION (LREC 2018), 2018, : 3925 - 3932