• DocumentCode
    2208696
  • Title

    Constraint Based Dimension Correlation and Distance Divergence for Clustering High-Dimensional Data

  • Author

    Zhang, Xianchao ; Wu, Yao ; Qiu, Yang

  • Author_Institution
    Sch. of Software, Dalian Univ. of Technol., Dalian, China
  • fYear
    2010
  • fDate
    13-17 Dec. 2010
  • Firstpage
    629
  • Lastpage
    638
  • Abstract
    Clusters are hidden in subspaces of high dimensional data, i.e., only a subset of features is relevant for each cluster. Subspace clustering is challenging since the search for the relevant features of each cluster and the detection of the final clusters are circular dependent and should be solved simultaneously. In this paper, we point out that feature correlation and distance divergence are important to subspace clustering, but both have not been considered in previous works. Feature correlation groups correlated features independently thus helps to reduce the search space for the relevant features search problem. Distance divergence distinguishes distances on different dimensions and helps to find the final clusters accurately. We tackle the two problems with the aid of a small amount domain knowledge in the form of must-links and cannot-links. We then devise a semi-supervised subspace clustering algorithm CDCDD. CDCDD integrates our solutions of the feature correlation and distance divergence problems, and uses an adaptive dimension voting scheme, which is derived from a previous unsupervised subspace clustering algorithm FINDIT. Experimental results on both synthetic data sets and real data sets show that the proposed CDCDD algorithm outperforms FINDIT in terms of accuracy, and outperforms the other constraint based algorithm SCMINER in terms of both accuracy and efficiency.
  • Keywords
    constraint handling; correlation methods; feature extraction; pattern clustering; search problems; set theory; CDCDD algorithm; adaptive dimension voting scheme; constraint based dimension correlation; distance divergence; feature correlation; feature search; high-dimensional data clustering; semisupervised subspace clustering algorithm; subset; high-dimensional data; pair-wise constraint; semi-supervised learning; subspace clustering;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Mining (ICDM), 2010 IEEE 10th International Conference on
  • Conference_Location
    Sydney, NSW
  • ISSN
    1550-4786
  • Print_ISBN
    978-1-4244-9131-5
  • Electronic_ISBN
    1550-4786
  • Type

    conf

  • DOI
    10.1109/ICDM.2010.15
  • Filename
    5694017