Title :
Text clustering approach based on maximal frequent term sets
Author :
Su, Chong ; Chen, Qingcai ; Wang, Xiaolong ; Meng, Xianjun
Author_Institution :
Intell. Comput. Res. Center, Harbin Inst. of Technol., Shenzhen, China
Abstract :
Classical text clustering algorithms are usually based on vector space model or its variants. Because of the high computing complexity and the difficulty of controlling clustering results, this kind of approaches are hard to be applied for the purpose of the large scale text clustering. Clustering algorithms based on frequent term sets make use of relationship among documents and their shared frequent term sets to achieve high accuracy and effectiveness in clustering. But since the number of frequent terms is usually too large to reach the efficiency requirement for large collection texts clustering, this paper proposes a novel text clustering approach based on maximal frequent term sets (MFTSC). This approach firstly mines maximal frequent term sets from text set and then clusters texts by following steps: at first, the maximal frequent term sets are clustered based on the criterion of k-mismatch; then texts are clustered according to term sets clustering results; finally, we categorize the left texts uncovered in previous step into produced text clusters Be compared with existing approaches, our experimental results show an average gain of 10% on F-Measure score with better performance on scalability and efficiency.
Keywords :
pattern clustering; set theory; text analysis; F-Measure score; k-mismatch; maximal frequent term sets; text clustering algorithm; text set; vector space model; Clustering algorithms; Cybernetics; Data mining; Large-scale systems; Motion pictures; Performance gain; Scalability; Space technology; Text mining; USA Councils; Frequent Term Sets; Maximal Frequent Term Sets; Text Clustering; Text mining;
Conference_Titel :
Systems, Man and Cybernetics, 2009. SMC 2009. IEEE International Conference on
Conference_Location :
San Antonio, TX
Print_ISBN :
978-1-4244-2793-2
Electronic_ISBN :
1062-922X
DOI :
10.1109/ICSMC.2009.5346313