Natural language processing for detecting cyberbullying in social media

Ramlan, Mohamad Radzi (2026) Natural language processing for detecting cyberbullying in social media. [Student Project] (Unpublished)
Abstract

Cyberbullying in Malaysia is a critical issue exacerbated by the widespread use of code-mixed Manglish on social media. Traditional keyword-based moderation systems failed to interpret the local slang and complex informal grammar, leading to high rates of undetected harassment. To address this technical gap, this study developed a localised detection framework by integrating a custom Manglish slang normalisation pipeline with a fine-tuned Bidirectional Encoder Representations from Transformers model. The model was trained on a dataset comprising 26,985 social media comments and deployed within an interactive web prototype equipped with explainable artificial intelligence features. The empirical evaluation demonstrated that the fine-tuned transformer model achieved an accuracy of 96.4% and an F1-score of 0.950. This performance significantly outperformed the traditional Support Vector Machine baseline, which only attained an 81.5% accuracy rate. Furthermore, the slang normalisation pipeline successfully reduced subword fragmentation and preserved contextual semantics before classification. The resulting web-based prototype processed inferences efficiently, providing transparent insights for human moderators. The findings established that combining rule-based lexical normalisation with advanced transformer architectures provided a demonstrated scalability potential and accurate solution for automated content moderation in code-mixed linguistic environments.

Item Details
Edit Item
Edit Item
Downloads & Files
[thumbnail of 147028.pdf]
Text
147028.pdf
Download (197kB)
Location & Physical Holdings
Digital Copy
Physical Location
  • Bilik Koleksi Akses Terhad | PTAR Kampus Samarahan 2, Sarawak
Item Status
Processing
Indexing & Metrics