Examining the Impact of Feature Selection Methods on Text Classification

Karaca, Mehmet Fatih; Bayir, Safak

Examining the Impact of Feature Selection Methods on Text Classification

Tarih

2017

Yazarlar

Karaca, Mehmet Fatih

Bayir, Safak

Yayıncı

Science & Information Sai Organization Ltd

Erişim Hakkı

info:eu-repo/semantics/closedAccess

Özet

Feature selection that aims to determine and select the distinctive terms representing a best document is one of the most important steps of classification. With the feature selection, dimension of document vectors are reduced and consequently duration of the process is shortened. In this study, feature selection methods were studied in terms of dimension reduction rates, classification success rates, and dimension reduction-classification success relation. As classifiers, kNN (k-Nearest Neighbors) and SVM (Support Vector Machines) were used. 5 standard (Odds Ratio-OR, Mutual Information-MI, Information Gain-IG, Chi-Square-CHI and Document Frequency-DF), 2 combined (Union of Feature Selections-UFS and Correlation of Union of Feature Selections-CUFS) and 1 new (Sum of Term Frequency-STF) feature selection methods were tested. The application was performed by selecting 100 to 1000 terms (with an increment of 100 terms) from each class. It was seen that kNN produces much better results than SVM. STF was found out to be the most successful feature selection considering the average values in both datasets. It was also found out that CUFS, a combined model, is the one that reduces the dimension the most, accordingly, it was seen that CUFS classify the documents more successfully with less terms and in short period compared to many of the standard methods.

Anahtar Kelimeler

Feature selection, text classification, text mining, k-Nearest Neighbors, support vector machines

Kaynak

International Journal of Advanced Computer Science and Applications

WoS Q Değeri

N/A

Cilt

8

Sayı

12

Bağlantı

https://hdl.handle.net/20.500.14619/8554

Koleksiyon

WoS İndeksli Yayınlar Koleksiyonu

Detaylı Öğe Kaydı

Examining the Impact of Feature Selection Methods on Text Classification

Tarih

Yazarlar

Dergi Başlığı

Dergi ISSN

Cilt Başlığı

Yayıncı

Erişim Hakkı

Özet

Açıklama

Anahtar Kelimeler

Kaynak

WoS Q Değeri

Scopus Q Değeri

Cilt

Sayı

Künye

Bağlantı

Koleksiyon