Improved Text Language Identification for the South African Languages

被引:0
|
作者
Duvenhage, Bernardt [1 ]
Ntini, Mfundo [2 ]
Ramonyai, Phala [2 ]
机构
[1] Praekelt Consulting, Feersun Engine, Johannesburg, South Africa
[2] Praekelt Consulting, Engn Team, Johannesburg, South Africa
来源
2017 PATTERN RECOGNITION ASSOCIATION OF SOUTH AFRICA AND ROBOTICS AND MECHATRONICS (PRASA-ROBMECH) | 2017年
关键词
Naive Bayesian text classification; lexicon based text classification; text language identification;
D O I
暂无
中图分类号
TP18 [人工智能理论];
学科分类号
081104 ; 0812 ; 0835 ; 1405 ;
摘要
Virtual assistants and text chatbots have recently been gaining popularity. Given the short message nature of text-based chat interactions, the language identification systems of these bots might only have 15 or 20 characters to make a prediction. However, accurate text language identification is important, especially in the early stages of many multilingual natural language processing pipelines. This paper investigates the use of a naive Bayes classifier, to accurately predict the language family that a piece of text belongs to, combined with a lexicon based classifier to distinguish the specific South African language that the text is written in. This approach leads to a 31% reduction in the language detection error. In the spirit of reproducible research the training and testing datasets as well as the code are published on github. Hopefully it will be useful to create a text language identification shared task for South African languages.
引用
收藏
页码:214 / 218
页数:5
相关论文
共 50 条
  • [1] Text-based language identification for South African languages
    Botha, Gerrit
    Zimu, Victor
    Barnard, Etienne
    SAIEE Africa Research Journal, 2007, 98 (04) : 141 - 148
  • [2] Language identification system for South African languages
    Mashao, DJ
    PROCEEDINGS OF THE 1998 SOUTH AFRICAN SYMPOSIUM ON COMMUNICATIONS AND SIGNAL PROCESSING: COMSIG '98, 1998, : 193 - 196
  • [3] DEVELOPMENT OF A SPOKEN LANGUAGE IDENTIFICATION SYSTEM FOR SOUTH AFRICAN LANGUAGES
    Peche, M.
    Davel, M. H.
    Barnard, E.
    SAIEE AFRICA RESEARCH JOURNAL, 2009, 100 (04): : 97 - 103
  • [4] Language Identification for South African Bantu Languages Using Rank Order Statistics
    Dube, Meluleki
    Suleman, Hussein
    DIGITAL LIBRARIES AT THE CROSSROADS OF DIGITAL INFORMATION FOR THE FUTURE, ICADL 2019, 2019, 11853 : 283 - 289
  • [5] Developing Text Resources for Ten South African Languages
    Eiselen, Roald
    Puttkammer, Martin J.
    LREC 2014 - NINTH INTERNATIONAL CONFERENCE ON LANGUAGE RESOURCES AND EVALUATION, 2014, : 3698 - 3703
  • [6] Text-based Language Identification for Some of the Under-resourced Languages of South Africa
    Sefara, Tshephisho Joseph
    Manamela, Madimetja Jonas
    Malatji, Promise Tshepiso
    2016 THIRD INTERNATIONAL CONFERENCE ON ADVANCES IN COMPUTING, COMMUNICATION AND ENGINEERING (ICACCE 2016), 2016, : 303 - 307
  • [7] Language policy incongruity and African languages in postapartheid South Africa
    Beukes, Anne-Marie
    LANGUAGE MATTERS, 2009, 40 (01) : 35 - 55
  • [8] Optical Character Recognition and text cleaning in the indigenous South African languages
    Prinsloo, Danie J.
    Taljard, Elsabe
    Goosen, Michelle
    STELLENBOSCH PAPERS IN LINGUISTICS PLUS-SPIL PLUS, 2022, 64 : 165 - 187
  • [9] Text Normalisation in Text-to-Speech Synthesis for South African Languages: Native Number Expansion
    Schlunz, Georg I.
    Dlamini, Nkosikhona
    Tshoane, Alfred
    Ramunyisi, Stan
    2017 PATTERN RECOGNITION ASSOCIATION OF SOUTH AFRICA AND ROBOTICS AND MECHATRONICS (PRASA-ROBMECH), 2017, : 230 - 235
  • [10] Orthographic measures of language distances between the official South African languages
    Zulu, P. N.
    Botha, G.
    Barnard, E.
    LITERATOR-JOURNAL OF LITERARY CRITICISM COMPARATIVE LINGUISTICS AND LITERARY STUDIES, 2008, 29 (01): : 185 - 204