Title of article
Proximity-based k-partitions clustering with ranking for document categorization and analysis
Author/Authors
Mei، نويسنده , , Jian-Ping and Chen، نويسنده , , Lihui، نويسنده ,
Issue Information
روزنامه با شماره پیاپی سال 2014
Pages
11
From page
7095
To page
7105
Abstract
As one of the most fundamental yet important methods of data clustering, center-based partitioning approach clusters the dataset into k subsets, each of which is represented by a centroid or medoid. In this paper, we propose a new medoid-based k-partitions approach called Clustering Around Weighted Prototypes (CAWP), which works with a similarity matrix. In CAWP, each cluster is characterized by multiple objects with different representative weights. With this new cluster representation scheme, CAWP aims to simultaneously produce clusters of improved quality and a set of ranked representative objects for each cluster. An efficient algorithm is derived to alternatingly update the clusters and the representative weights of objects with respect to each cluster. An annealing-like optimization procedure is incorporated to alleviate the local optimum problem for better clustering results and at the same time to make the algorithm less sensitive to parameter setting. Experimental results on benchmark document datasets show that, CAWP achieves favorable effectiveness and efficiency in clustering, and also provides useful information for cluster-specified analysis.
Keywords
Document categorization , Clustering , k-Medoids , Similarity-based , Partitioning
Journal title
Expert Systems with Applications
Serial Year
2014
Journal title
Expert Systems with Applications
Record number
2355204
Link To Document