مرکز منطقه ای اطلاع رساني علوم و فناوري - The <formula formulatype="inline"> <img src="/images/tex/523.gif" alt="K"> </formula>-Means-Type Algorithms Versus Imbalanced Data Distributions

DocumentCode :

1413313

Title :

The $K$ -Means-Type Algorithms Versus Imbalanced Data Distributions

Author :

Liang, Jiye ; Bai, Liang ; Dang, Chuangyin ; Cao, Fuyuan

Author_Institution :

Key Lab. of Comput. Intell. & Chinese Inf. Process. of Minist. of Educ., Shanxi Univ., Taiyuan, China

Volume :

Issue :

fYear :

2012

Firstpage :

728

Lastpage :

745

Abstract :

K-means is a partitional clustering technique that is well-known and widely used for its low computational cost. The representative algorithms include the hard k-means and the fuzzy k -means. However, the performance of these algorithms tends to be affected by skewed data distributions, i.e., imbalanced data. They often produce clusters of relatively uniform sizes, even if input data have varied cluster sizes, which is called the “uniform effect.” In this paper, we analyze the causes of this effect and illustrate that it probably occurs more in the fuzzy k-means clustering process than the hard k-means clustering process. As the fuzzy index m increases, the “uniform effect” becomes evident. To prevent the effect of the “uniform effect,” we propose a multicenter clustering algorithm in which multicenters are used to represent each cluster, instead of one single center. The proposed algorithm consists of the three subalgorithms: the fast global fuzzy k -means, Best M-Plot, and grouping multicenter algorithms. They will be, respectively, used to address the three important problems: 1) How are the reliable cluster centers from a dataset obtained? 2) How are the number of clusters which these obtained cluster centers represent determined? 3) How is it judged as to which cluster centers represent the same clusters? The experimental studies on both synthetic and real datasets illustrate the effectiveness of the proposed clustering algorithm in clustering balanced and imbalanced data.

Keywords :

fuzzy set theory; pattern clustering; best m-plot; fast global fuzzy k-means; fuzzy index; fuzzy k-means clustering process; grouping multicenter algorithms; hard k-means; imbalanced data distributions; k-means-type algorithms; multicenter clustering algorithm; partitional clustering technique; representative algorithms; skewed data distributions; uniform effect; Algorithm design and analysis; Clustering algorithms; Indexes; Kernel; Partitioning algorithms; Search problems; Shape; Imbalanced data; multirepresentatives; the $k$-means-type clustering algorithms; the number of clusters; the production of cluster centers;

fLanguage :

English

Journal_Title :

Fuzzy Systems, IEEE Transactions on

Publisher :

ieee

ISSN :

1063-6706

Type :

jour

DOI :

10.1109/TFUZZ.2011.2182354

Filename :

6121900

Link To Document :

https://search.ricest.ac.ir/dl/search/defaultta.aspx?DTC=49&DC=1413313