Title of article
A subspace decision cluster classifier for text classification
Author/Authors
Li، نويسنده , , Yan and Hung، نويسنده , , Edward Ho Chung Wong، نويسنده , , Korris، نويسنده ,
Issue Information
روزنامه با شماره پیاپی سال 2011
Pages
8
From page
12475
To page
12482
Abstract
In this paper, a new classification method (SDCC) for high dimensional text data with multiple classes is proposed. In this method, a subspace decision cluster classification (SDCC) model consists of a set of disjoint subspace decision clusters, each labeled with a dominant class to determine the class of new objects falling in the cluster. A cluster tree is first generated from a training data set by recursively calling a subspace clustering algorithm Entropy Weighting k-Means algorithm. Then, the SDCC model is extracted from the subspace decision cluster tree. Various tests including Anderson–Darling test are used to determine the stopping condition of the tree growing. A series of experiments on real text data sets have been conducted. Their results show that the new classification method (SDCC) outperforms the existing methods like decision tree and SVM. SDCC is particularly suitable for large, high dimensional sparse text data with many classes.
Keywords
Classification , Subspace decision cluster
Journal title
Expert Systems with Applications
Serial Year
2011
Journal title
Expert Systems with Applications
Record number
2350255
Link To Document