A generalized cluster centroid based classifier for text categorization

► We focus on using a constrained clustering algorithm to integrate the KNN and Rocchio classifiers. ► Clustering can be used to strengthen the Rocchio model. ► The KNN categorization process can be greatly accelerated by using the improved Rocchio model together with the KNN decision rule to perfor...

Full description

Saved in:
Bibliographic Details
Published inInformation processing & management Vol. 49; no. 2; pp. 576 - 586
Main Authors Pang, Guansong, Jiang, Shengyi
Format Journal Article
LanguageEnglish
Published Kidlington Elsevier Ltd 01.03.2013
Elsevier
Elsevier Science Ltd
Subjects
Online AccessGet full text
ISSN0306-4573
1873-5371
DOI10.1016/j.ipm.2012.10.003

Cover

More Information
Summary:► We focus on using a constrained clustering algorithm to integrate the KNN and Rocchio classifiers. ► Clustering can be used to strengthen the Rocchio model. ► The KNN categorization process can be greatly accelerated by using the improved Rocchio model together with the KNN decision rule to perform categorization. ► Extensive experiments on heterogeneous corpora show the effectiveness and efficiency of our proposed method. In this paper, a Generalized Cluster Centroid based Classifier (GCCC) and its variants for text categorization are proposed by utilizing a clustering algorithm to integrate two well-known classifiers, i.e., the K-nearest-neighbor (KNN) classifier and the Rocchio classifier. KNN, a lazy learning method, suffers from inefficiency in online categorization while achieving remarkable effectiveness. Rocchio, which has efficient categorization performance, fails to obtain an expressive categorization model due to its inherent linear separability assumption. Our proposed method mainly focuses on two points: one point is that we use a clustering algorithm to strengthen the expressiveness of the Rocchio model; another one is that we employ the improved Rocchio model to speed up the categorization process of KNN. Extensive experiments conducted on both English and Chinese corpora show that GCCC and its variants have better categorization ability than some state-of-the-art classifiers, i.e., Rocchio, KNN and Support Vector Machine (SVM).
Bibliography:SourceType-Scholarly Journals-1
ObjectType-Feature-1
content type line 14
ObjectType-Article-2
content type line 23
ObjectType-Article-1
ObjectType-Feature-2
ISSN:0306-4573
1873-5371
DOI:10.1016/j.ipm.2012.10.003