Principal Component Analysis (PCA) is a technique that reduces data dimensionality by converting correlated variables into uncorrelated components. It simplifies data while maintaining patterns. Combining PCA with Random Forest (RF) can increase efficiency and reduce overfitting. Nonetheless, its benefits depend on data standardization. PCA is a statistical strategy that calculates the covariance matrix of a data set, extracting eigenvectors and eigenvalues to minimize dimensionality while preserving significant variance. It involves covariance-based eigenvalue decomposition and Singular Value Decomposition (SVD), with the best PCs accounting for 95% of the variance. Modified PCA data can be used for training RF, but potential signal loss and interpretability issues may arise. High-dimensional data classification in genomics and medical diagnostics faces challenges due to noise accumulation, class imbalance, and dimensionality. Biased classification, noise buildup, and instability in feature selection hinder performance. Misclassification of training data affects both supervised and unsupervised techniques, leading to increased error rates and resource inefficiency. RF has garnered considerable attention from various researchers among numerous classification approaches. However, RF performs poorly with very high- dimensional data. Therefore, the objective of this study was to classify high- dimensional data using RF alone and a combined RF with PCA and Sparse PCA. The study used SVD and Random Orthogonal Matrix (ROM) to reduce a high-dimensional dataset. The Variance Maximization (VM), Reconstruction Error Minimization (REM), and SVD are effective methods for high-dimensional datasets, enhancing data variability assessment, model resilience, and dimensionality reduction, facilitating efficient summarization, noise reduction, and data analysis or predictive modeling. A PCA strategy that minimizes data reconstruction error, ensuring condensed data faithfully depicts original data, is instrumental in image processing and other critical data reconstruction and compression applications. SVD is a factorization that decomposes a matrix into three components: U, D, and V, which are used for tasks such as dimensionality reduction and solving systems of equations. The results show that RF alone (accuracy = 93.98%) performed best compared with VM combined with RF (accuracy = 73.23%), RF combined with PCA (accuracy = 92.48%), REM with RF (93.33%) and SVD with RF (93.33%). Sparse PCA is an alternative method for addressing dimensionality issues that hinder classification performance in RFs, particularly in genomic studies. The findings highlight that Sparse PCA techniques, particularly, effectively address dimensionality issues and significantly enhance classification performance in RF, especially for genomic applications. The study indicates that integrating dimensionality-reduction methods with RF classification substantially improves performance in high-dimensional domains, such as genomics. Techniques such as PCA, Sparse PCA,
| Item Type: | Thesis (Masters) |
|---|---|
| Creators: | Creators Email / ID Num. Ahmed Jasim, Zahwa 2021740989 |
| Contributors: | Contribution Name Email / ID Num. Advisor Muhammad Japeri, Ahmed Zia ul- Saufie UNSPECIFIED |
| Subjects: | Q Science > QA Mathematics Q Science > QA Mathematics > Multivariate analysis. Cluster analysis. Longitudinal method |
| Divisions: | Universiti Teknologi MARA, Shah Alam > Faculty of Computer and Mathematical Sciences |
| Programme: | Master of Science (Statistics) |
| Keywords: | Data welding, High dimension, Principal Component Analysis (PCA) |
| Date: | June 2026 |
| URI: | https://ir.uitm.edu.my/id/eprint/145654 |
145654_fulltext.pdf
Available under License Dasar Harta Intelek UiTM (Para 6).
declarationform.pdf
Restricted to Repository staff only

