Abstract
Social media data presents significant challenges for natural language processing because they frequently contain non-standard language forms, including creative spellings, shortened expressions, phonetic variants, mixed-language usage, and informal constructions. While existing research has extensively explored text normalization techniques, far less attention has been given to how such noisy elements are identified and categorized before normalization takes place. This scoping review examines the approaches used to detect and classify noisy text prior to normalization. A systematic search was carried out across Web of Science, Scopus, IEEE Xplore, ACM Digital Library, and ScienceDirect, covering studies published between 2020 and 2025. Following duplicate removal and screening, 35 studies were retained from an initial pool of 770 records. The review addresses three main questions concerning commonly used identification techniques, defining characteristics of noisy text, and challenges encountered during normalization. The findings indicate that researchers employ diverse strategies ranging from rule-based and dictionary-based methods to machine learning, deep learning, and similarity-based approaches. Common noise characteristics include spelling variation, phonetic alteration, code-mixing, and informal word forms. Persistent challenges relate to the absence of shared benchmarks, limited annotated and multilingual resources, and difficulties in applying detection methods across different domains. By explicitly focusing on the identification stage, this review clarifies an often-implicit step in text normalization pipelines and offers direction for future work, particularly in low-resource and informal language settings.
Metadata
| Item Type: | Article |
|---|---|
| Creators: | Creators Email / ID Num. Ahmad, Helmi Ashraf helmiashraf10@gmail.com Azizan, Azilawati azila899@uitm.edu.my Mohamed Hanum, Haslizatul UNSPECIFIED |
| Subjects: | P Language and Literature > P Philology. Linguistics > Study and teaching. Research > Computer-assisted instruction Q Science > QA Mathematics > Instruments and machines > Electronic Computers. Computer Science |
| Divisions: | Universiti Teknologi MARA, Perak > Tapah Campus > Faculty of Computer and Mathematical Sciences |
| Journal or Publication Title: | Mathematical Sciences and Informatics Journal (MIJ) |
| UiTM Journal Collections: | UiTM Journals > Mathematical Science and Information Journal (MIJ) |
| ISSN: | 2735-0703 |
| Volume: | 7 |
| Number: | 1 |
| Page Range: | pp. 59-83 |
| Keywords: | Scoping review, Noisy text, Text normalization, Natural language processing, Text analysis |
| Date: | April 2026 |
| URI: | https://ir.uitm.edu.my/id/eprint/141729 |
