Identifying noisy Malay and Malay-English code-mixed social media text before normalization: a scoping review

Ahmad, Helmi Ashraf and Azizan, Azilawati and Mohamed Hanum, Haslizatul (2026) Identifying noisy Malay and Malay-English code-mixed social media text before normalization: a scoping review. Mathematical Sciences and Informatics Journal (MIJ), 7 (1). pp. 59-83. ISSN 2735-0703

Official URL: https://mijuitm.com.my

Identification Number (DOI): 10.24191/mij.v7i1.11574

Abstract

Social media data presents significant challenges for natural language processing because they frequently contain non-standard language forms, including creative spellings, shortened expressions, phonetic variants, mixed-language usage, and informal constructions. While existing research has extensively explored text normalization techniques, far less attention has been given to how such noisy elements are identified and categorized before normalization takes place. This scoping review examines the approaches used to detect and classify noisy text prior to normalization. A systematic search was carried out across Web of Science, Scopus, IEEE Xplore, ACM Digital Library, and ScienceDirect, covering studies published between 2020 and 2025. Following duplicate removal and screening, 35 studies were retained from an initial pool of 770 records. The review addresses three main questions concerning commonly used identification techniques, defining characteristics of noisy text, and challenges encountered during normalization. The findings indicate that researchers employ diverse strategies ranging from rule-based and dictionary-based methods to machine learning, deep learning, and similarity-based approaches. Common noise characteristics include spelling variation, phonetic alteration, code-mixing, and informal word forms. Persistent challenges relate to the absence of shared benchmarks, limited annotated and multilingual resources, and difficulties in applying detection methods across different domains. By explicitly focusing on the identification stage, this review clarifies an often-implicit step in text normalization pipelines and offers direction for future work, particularly in low-resource and informal language settings.

Metadata

Item Type: Article
Creators:
Creators
Email / ID Num.
Ahmad, Helmi Ashraf
helmiashraf10@gmail.com
Azizan, Azilawati
azila899@uitm.edu.my
Mohamed Hanum, Haslizatul
UNSPECIFIED
Subjects: P Language and Literature > P Philology. Linguistics > Study and teaching. Research > Computer-assisted instruction
Q Science > QA Mathematics > Instruments and machines > Electronic Computers. Computer Science
Divisions: Universiti Teknologi MARA, Perak > Tapah Campus > Faculty of Computer and Mathematical Sciences
Journal or Publication Title: Mathematical Sciences and Informatics Journal (MIJ)
UiTM Journal Collections: UiTM Journals > Mathematical Science and Information Journal (MIJ)
ISSN: 2735-0703
Volume: 7
Number: 1
Page Range: pp. 59-83
Keywords: Scoping review, Noisy text, Text normalization, Natural language processing, Text analysis
Date: April 2026
URI: https://ir.uitm.edu.my/id/eprint/141729
Edit Item
Edit Item

Download

[thumbnail of 141729.pdf] Text
141729.pdf

Download (505kB)

ID Number

141729

Indexing

Altmetric
PlumX
Dimensions

Statistic

Statistic details