Multilingual Hate Speech Detection: A Semi-Supervised Generative Adversarial Approach

被引:1
|
作者
Mnassri, Khouloud [1 ]
Farahbakhsh, Reza [1 ]
Crespi, Noel [1 ]
机构
[1] Inst Polytech Paris, Samovar Telecom SudParis, F-91120 Palaiseau, France
关键词
social media; hate speech; semisupervised; GAN; multilingual; PLMs; DATA AUGMENTATION; NETWORKS;
D O I
10.3390/e26040344
中图分类号
O4 [物理学];
学科分类号
0702 ;
摘要
Social media platforms have surpassed cultural and linguistic boundaries, thus enabling online communication worldwide. However, the expanded use of various languages has intensified the challenge of online detection of hate speech content. Despite the release of multiple Natural Language Processing (NLP) solutions implementing cutting-edge machine learning techniques, the scarcity of data, especially labeled data, remains a considerable obstacle, which further requires the use of semisupervised approaches along with Generative Artificial Intelligence (Generative AI) techniques. This paper introduces an innovative approach, a multilingual semisupervised model combining Generative Adversarial Networks (GANs) and Pretrained Language Models (PLMs), more precisely mBERT and XLM-RoBERTa. Our approach proves its effectiveness in the detection of hate speech and offensive language in Indo-European languages (in English, German, and Hindi) when employing only 20% annotated data from the HASOC2019 dataset, thereby presenting significantly high performances in each of multilingual, zero-shot crosslingual, and monolingual training scenarios. Our study provides a robust mBERT-based semisupervised GAN model (SS-GAN-mBERT) that outperformed the XLM-RoBERTa-based model (SS-GAN-XLM) and reached an average F1 score boost of 9.23% and an accuracy increase of 5.75% over the baseline semisupervised mBERT model.
引用
收藏
页数:19
相关论文
共 50 条
  • [1] Multilingual Hate Speech Detection Using Semi-supervised Generative Adversarial Network
    Mnassri, Khouloud
    Farahbakhsh, Reza
    Crespi, Noel
    [J]. COMPLEX NETWORKS & THEIR APPLICATIONS XII, VOL 4, COMPLEX NETWORKS 2023, 2024, 1144 : 192 - 204
  • [2] Semi-Supervised Learning with Generative Adversarial Networks for Pathological Speech Classification
    Trinh, Nam H.
    O'Brien, Darragh
    [J]. 2020 31ST IRISH SIGNALS AND SYSTEMS CONFERENCE (ISSC), 2020, : 214 - 218
  • [3] Generative Adversarial Training for Supervised and Semi-supervised Learning
    Wang, Xianmin
    Li, Jing
    Liu, Qi
    Zhao, Wenpeng
    Li, Zuoyong
    Wang, Wenhao
    [J]. FRONTIERS IN NEUROROBOTICS, 2021, 15
  • [4] SEMI-SUPERVISED CHANGE DETECTION BASED ON GRAPHS WITH GENERATIVE ADVERSARIAL NETWORKS
    Liu, Junfu
    Chen, Keming
    Xu, Guangluan
    Li, Hao
    Yan, Menglong
    Diao, Wenhui
    Sun, Xian
    [J]. 2019 IEEE INTERNATIONAL GEOSCIENCE AND REMOTE SENSING SYMPOSIUM (IGARSS 2019), 2019, : 74 - 77
  • [5] Semi-supervised community detection method based on generative adversarial networks
    Liu, Xiaoyang
    Zhang, Mengyao
    Liu, Yanfei
    Liu, Chao
    Li, Chaorong
    Wang, Wei
    Zhang, Xiaoqin
    Bouyer, Asgarali
    [J]. JOURNAL OF KING SAUD UNIVERSITY-COMPUTER AND INFORMATION SCIENCES, 2024, 36 (03)
  • [6] DISCRIMINATIVE SEMI-SUPERVISED GENERATIVE ADVERSARIAL NETWORK FOR HYPERSPECTRAL ANOMALY DETECTION
    Jiang, Tao
    Xie, Weiying
    Li, Yunsong
    Du, Qian
    [J]. IGARSS 2020 - 2020 IEEE INTERNATIONAL GEOSCIENCE AND REMOTE SENSING SYMPOSIUM, 2020, : 2420 - 2423
  • [7] Semi-Supervised Self-Learning for Arabic Hate Speech Detection
    Alsafari, Safa
    Sadaoui, Samira
    [J]. 2021 IEEE INTERNATIONAL CONFERENCE ON SYSTEMS, MAN, AND CYBERNETICS (SMC), 2021, : 863 - 868
  • [8] Poster Abstract: A Semi-Supervised Approach for Network Intrusion Detection Using Generative Adversarial Networks
    Jeong, Hyejeong
    Yu, Jieun
    Lee, Wonjun
    [J]. IEEE CONFERENCE ON COMPUTER COMMUNICATIONS WORKSHOPS (IEEE INFOCOM WKSHPS 2021), 2021,
  • [9] Generative adversarial network for semi-supervised image captioning
    Liang, Xu
    Li, Chen
    Tian, Lihua
    [J]. Computer Vision and Image Understanding, 2024, 249
  • [10] Semi-supervised Seizure Prediction with Generative Adversarial Networks
    Nhan Duy Truong
    Zhou, Luping
    Kavehei, Omid
    [J]. 2019 41ST ANNUAL INTERNATIONAL CONFERENCE OF THE IEEE ENGINEERING IN MEDICINE AND BIOLOGY SOCIETY (EMBC), 2019, : 2369 - 2372