Exhaustive Study into Machine Learning and Deep Learning Methods for Multilingual Cyberbullying Detection in Bangla and Chittagonian Texts

被引：5

作者：

Mahmud, Tanjim ^{[1
,2
]}

Ptaszynski, Michal ^{[1
]}

Masui, Fumito ^{[1
]}

机构：

[1] Kitami Inst Technol, Text Informat Proc Lab, 165 Koen Cho, Kitami City, Hokkaido 0908507, Japan

[2] Rangamati Sci & Technol Univ, Dept Comp Sci & Engn, Rangamati 4500, Bangladesh

来源：

ELECTRONICS | 2024年 / 13卷 / 09期

关键词：

multilingual models; low-resource languages; machine learning; ensemble models; deep learning; hybrid models; transformers models; MODEL;

D O I：

10.3390/electronics13091677

中图分类号：

TP [自动化技术、计算机技术];

学科分类号：

0812 ;

摘要：

Cyberbullying is a serious problem in online communication. It is important to find effective ways to detect cyberbullying content to make online environments safer. In this paper, we investigated the identification of cyberbullying contents from the Bangla and Chittagonian languages, which are both low-resource languages, with the latter being an extremely low-resource language. In the study, we used both traditional baseline machine learning methods, as well as a wide suite of deep learning methods especially focusing on hybrid networks and transformer-based multilingual models. For the data, we collected over 5000 both Bangla and Chittagonian text samples from social media. Krippendorff's alpha and Cohen's kappa were used to measure the reliability of the dataset annotations. Traditional machine learning methods used in this research achieved accuracies ranging from 0.63 to 0.711, with SVM emerging as the top performer. Furthermore, employing ensemble models such as Bagging with 0.70 accuracy, Boosting with 0.69 accuracy, and Voting with 0.72 accuracy yielded promising results. In contrast, deep learning models, notably CNN, achieved accuracies ranging from 0.69 to 0.811, thus outperforming traditional ML approaches, with CNN exhibiting the highest accuracy. We also proposed a series of hybrid network-based models, including BiLSTM+GRU with an accuracy of 0.799, CNN+LSTM with 0.801 accuracy, CNN+BiLSTM with 0.78 accuracy, and CNN+GRU with 0.804 accuracy. Notably, the most complex model, (CNN+LSTM)+BiLSTM, attained an accuracy of 0.82, thus showcasing the efficacy of hybrid architectures. Furthermore, we explored transformer-based models, such as XLM-Roberta with 0.841 accuracy, Bangla BERT with 0.822 accuracy, Multilingual BERT with 0.821 accuracy, BERT with 0.82 accuracy, and Bangla ELECTRA with 0.785 accuracy, which showed significantly enhanced accuracy levels. Our analysis demonstrates that deep learning methods can be highly effective in addressing the pervasive issue of cyberbullying in several different linguistic contexts. We show that transformer models can efficiently circumvent the language dependence problem that plagues conventional transfer learning methods. Our findings suggest that hybrid approaches and transformer-based embeddings can effectively tackle the problem of cyberbullying across online platforms.

引用

页数：36

共 50 条

[41] A Review on Deep-Learning-Based Cyberbullying Detection
Hasan, Md. Tarek
Hossain, Md. Al Emran
Mukta, Md. Saddam Hossain
Akter, Arifa
Ahmed, Mohiuddin
Islam, Salekul
FUTURE INTERNET, 2023, 15 (05)
[42] CYBERBULLYING DETECTION ON TIKTOK USING A DEEP LEARNING APPROACH
Stoleriu, Razvan
Nascu, Andrei
Pop, Florin
UNIVERSITY POLITEHNICA OF BUCHAREST SCIENTIFIC BULLETIN SERIES C-ELECTRICAL ENGINEERING AND COMPUTER SCIENCE, 2025, 87 (01): : 5 - 20
[43] Cyberbullying detection solutions based on deep learning architectures
Celestine Iwendi
Gautam Srivastava
Suleman Khan
Praveen Kumar Reddy Maddikunta
Multimedia Systems, 2023, 29 : 1839 - 1852
[44] Cyberbullying Detection in Social Networks Using Deep Learning
Khafajeh, Hayel
INTERNATIONAL ARAB JOURNAL OF INFORMATION TECHNOLOGY, 2024, 21 (06) : 1054 - 1063
[45] Cyberbullying detection solutions based on deep learning architectures
Iwendi, Celestine
Srivastava, Gautam
Khan, Suleman
Maddikunta, Praveen Kumar Reddy
MULTIMEDIA SYSTEMS, 2023, 29 (03) : 1839 - 1852
[46] Cyberbullying detection from tweets using deep learning
Bharti, Shubham
Yadav, Arun Kumar
Kumar, Mohit
Yadav, Divakar
KYBERNETES, 2022, 51 (09) : 2695 - 2711
[47] Optimized Twitter Cyberbullying Detection based on Deep Learning
Al-Ajlan, Monirah A.
Ykhlef, Mourad
2018 21ST SAUDI COMPUTER SOCIETY NATIONAL COMPUTER CONFERENCE (NCC), 2018,
[48] A Study: Machine Learning and Deep Learning Approaches for Intrusion Detection System
Sekhar, C. H.
Rao, K. Venkata
SECOND INTERNATIONAL CONFERENCE ON COMPUTER NETWORKS AND COMMUNICATION TECHNOLOGIES, ICCNCT 2019, 2020, 44 : 845 - 849
[49] Classifying Multilingual User Feedback using Traditional Machine Learning and Deep Learning
Stanik, Christoph
Haering, Marlo
Maalej, Walid
2019 IEEE 27TH INTERNATIONAL REQUIREMENTS ENGINEERING CONFERENCE WORKSHOPS (REW 2019), 2019, : 220 - 226
[50] Cyberbullying detection through deep learning: A case study of Turkish celebrities on Twitter
Karadag, Bulut
Akbulut, Akhan
Zaim, Abdul Halim
WEB INTELLIGENCE, 2023, 21 (01) : 61 - 70

← 1 2 3 4 5 →