• Title of article

    A new initialization method for categorical data clustering

  • Author/Authors

    Cao، نويسنده , , Fuyuan and Liang، نويسنده , , Jiye and Bai، نويسنده , , Liang، نويسنده ,

  • Issue Information
    روزنامه با شماره پیاپی سال 2009
  • Pages
    6
  • From page
    10223
  • To page
    10228
  • Abstract
    In clustering algorithms, choosing a subset of representative examples is very important in data set. Such “exemplars” can be found by randomly choosing an initial subset of data objects and then iteratively refining it, but this works well only if that initial choice is close to a good solution. In this paper, based on the frequency of attribute values, the average density of an object is defined. Furthermore, a novel initialization method for categorical data is proposed, in which the distance between objects and the density of the object is considered. We also apply the proposed initialization method to k-modes algorithm and fuzzy k-modes algorithm. Experimental results illustrate that the proposed initialization method is superior to random initialization method and can be applied to large data sets for its linear time complexity with respect to the number of data objects.
  • Keywords
    Initial cluster center , k-modes algorithm , Initialization method , distance , Density
  • Journal title
    Expert Systems with Applications
  • Serial Year
    2009
  • Journal title
    Expert Systems with Applications
  • Record number

    2346785