Text categorization with support vector machines

Khanafi, Nur Amira (2013) Text categorization with support vector machines. [Student Project] (Unpublished)
Abstract

With the growing number of text documents on the internet made it difficult for users to search, find, manage and organize information quickly. In the past, text documents are classified manually and it is time-consuming. Text categorization is a process of assigning text documents into a set of fixed predefined categories. The high dimensionality of text documents made it difficult to categorize because text documents contain noise and useless data. This project explained several methods of feature selection that can be used to reduce high dimensionality of feature space in text documents such as Information Gain, Gain Ratio, CHI-Squares, Mutual Information and Document frequency. This project also explained how text categorization is developed using Support Vector Machines and solved the text categorization problems. The results showed that Support Vector Machines perform well and very fast at both training and testing. It also showed that Support Vector Machines is a suitable technique to handle text categorization since this project are using large amount of datasets.

Item Details
Edit Item
Edit Item
Downloads & Files
[thumbnail of 136461.pdf]
Text
136461.pdf
Download (907kB)
Location & Physical Holdings
Digital Copy
Physical Location
Item Status
Indexing & Metrics
Download Statistics