• DocumentCode
    1018704
  • Title

    Text Clustering with Feature Selection by Using Statistical Data

  • Author

    Li, Yanjun ; Luo, Congnan ; Chung, Soon M.

  • Author_Institution
    Fordham Univ., Bronx
  • Volume
    20
  • Issue
    5
  • fYear
    2008
  • fDate
    5/1/2008 12:00:00 AM
  • Firstpage
    641
  • Lastpage
    652
  • Abstract
    Feature selection is an important method for improving the efficiency and accuracy of text categorization algorithms by removing redundant and irrelevant terms from the corpus. In this paper, we propose a new supervised feature selection method, named CHIR, which is based on the chi2 statistic and new statistical data that can measure the positive term-category dependency. We also propose a new text clustering algorithm, named text clustering with feature selection (TCFS). TCFS can incorporate CHIR to identify relevant features (i.e., terms) iteratively, and the clustering becomes a learning process. We compared TCFS and the K-means clustering algorithm in combination with different feature selection methods for various real data sets. Our experimental results show that TCFS with CHIR has better clustering accuracy in terms of the F-measure and the purity.
  • Keywords
    pattern clustering; statistical analysis; text analysis; CHIR; K-means clustering; chi2 statistic; statistical data; supervised feature selection; text categorization; text clustering; Chi-square statistics; Feature selection; Performance analysis; Text clustering; Text mining;
  • fLanguage
    English
  • Journal_Title
    Knowledge and Data Engineering, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1041-4347
  • Type

    jour

  • DOI
    10.1109/TKDE.2007.190740
  • Filename
    4408578