Multiple Linear Regression and K-Mean Clustering of post-covid on household income in Malaysia

Mohd Sabri, Muhammad Akmal Danial and Wan Shahidan, Wan Nurshazelin (2025) Multiple Linear Regression and K-Mean Clustering of post-covid on household income in Malaysia. In: Proceedings of Research Exhibition in Mathematics and Computer Sciences 2025 (REMACS 8.0). Faculty of Computer and Mathematical Sciences, UiTM Cawangan Perlis, pp. 9-10.
Abstract

This study explores the socio-economic factors that influence household income in Malaysia after the COVID-19 pandemic by using a combination of Multiple Linear Regression and K-Means clustering. Data were taken from the 2022 Household Income and Expenditure Survey provided by the Department of Statistics Malaysia. The regression analysis found that household size, age, gender, and education level significantly affect income. The model achieved an adjusted R-squared value of 0.499 and an F-value of 35.727, showing a moderately strong fit. These significant variables were then used in the clustering process, which grouped households into three distinct segments. The first group includes older, low-income, and less-educated households. The second group represents middleincome households with moderate education. The third group consists of younger, highly educated, and higher-income households. This hybrid approach gave more detailed insights than using regression alone and supports more targeted and effective policy planning. The findings are consistent with earlier research by Yee et al. in 2023, which highlighted the importance of clustering in understanding income differences in Malaysia.

Item Details
Edit Item
Edit Item
Downloads & Files
[thumbnail of 143689.pdf]
Text
143689.pdf
Download (248kB)
Indexing & Metrics