Deep Learning Approaches for Hate Speech Detection in Resource-Constrained Social Media Texts: A Comparative Evaluation

Bastian, Yosia Immanuel and Hermawan, Aditiya and Maranto, Ardiane Rossi Kurniawan and Daniawan, Benny and Junaedi, Junaedi (2026) Deep Learning Approaches for Hate Speech Detection in Resource-Constrained Social Media Texts: A Comparative Evaluation. Artificial Intelligence and Applications, 00 (00). 01-11. ISSN 2811-0854

[img] Text
1 Jurnal Q1 AIA Benny Deep Learning Approaches for Hate Speech.pdf - Published Version

Download (837kB)
[img] Text
1 Turnitin Q1 AIA Benny Deep Learning Approaches for Hate Speech.pdf - Published Version

Download (3MB)

Abstract

Hate speech that occurred in social media platforms presents critical safety challenges, especially in language environments with minimal source where text is often written in an informal form, abbreviated, and context dependent. This paper further provides empirical evidence regarding the ability of the BERT (Bidirectional Encoder Representations from Transformers) model in detecting hate speech in a rough text environment and comparing directly with other deep learning techniques.Three Bidirectional Long ShortTerm Memory (BiLSTM) variants using Word2Vec, FastText, and TF-IDF representations were compared with BERT, which employs contextual language representations. The data is separated by using a train-test ratio of 80:20, and the performance is evaluated according to their accuracy, precision, recall, F1-score, and area under the curve (AUC). The result shows that BERT outperforms all variations of BiLSTM by achieving an accuracy of 87.18%, an F1-score of 87.08%, and an AUC matrix of 0.9244. According to the efficiency matters, BiLSTM and Term Frequency–Inverse Document Frequency (TF-IDF) are the most ineffective for their uneven classification distribution, while BiLSTM with Word2Vec and FastText shows moderate effectivity. This finding conclusively demonstrates the benefit of transformer-based models to capture the subtleties of language with noisy textual data, which is commonly found in low-resource settings. To conclude the above discussion, this finding suggests that BERT has great potential as a multilingual content moderation tool to be applied in informal and unstructured digital environments.

Item Type: Article
Uncontrolled Keywords: BERT, deep learning, hate speech detection, low-resource language, BiLSTM
Subjects: 000 Karya Umum > 003 Sistem-sistem > 003.3 Model dan Simulasi Komputer
Divisions: Fakultas Sains & Teknologi > Sistem Informasi
Depositing User: Muhamad Kemal Prasetyo
Date Deposited: 31 Aug 2026 04:28
Last Modified: 31 Aug 2026 04:28
URI: https://repositori.buddhidharma.ac.id//id/eprint/3494

Actions (login required)

View Item View Item