With the growing number of text documents on the internet made it difficult for users to search, find, manage and organize information quickly. In the past, text documents are classified manually and it is time-consuming. Text categorization is a process of assigning text documents into a set of fixed predefined categories. The high dimensionality of text documents made it difficult to categorize because text documents contain noise and useless data. This project explained several methods of feature selection that can be used to reduce high dimensionality of feature space in text documents such as Information Gain, Gain Ratio, CHI-Squares, Mutual Information and Document frequency. This project also explained how text categorization is developed using Support Vector Machines and solved the text categorization problems. The results showed that Support Vector Machines perform well and very fast at both training and testing. It also showed that Support Vector Machines is a suitable technique to handle text categorization since this project are using large amount of datasets.
| Item Type: | Student Project |
|---|---|
| Creators: | Creators Email / ID Num. Khanafi, Nur Amira UNSPECIFIED |
| Contributors: | Contribution Name Email / ID Num. Advisor Abdul Rahman, Shuzlina UNSPECIFIED |
| Subjects: | Q Science > Q Science (General) Q Science > QA Mathematics |
| Divisions: | Universiti Teknologi MARA, Shah Alam > Faculty of Computer and Mathematical Sciences |
| Programme: | Bachelor of Science (Hons) Intelligent System |
| Keywords: | Text categorization, Support vector machines, Feature selection methods |
| Date: | 2013 |
| URI: | https://ir.uitm.edu.my/id/eprint/136461 |
136461.pdf

