DocumentCode
1878326
Title
Mining Approximate Frequency Itemsets over Data Streams Based on D-Hash Table
Author
Ju, Chunhua ; You, Gang
Author_Institution
Coll. of Comput. Sci. & Inf. Eng., Zhejiang Gongshang Univ., Hangzhou, China
fYear
2009
fDate
27-29 May 2009
Firstpage
249
Lastpage
254
Abstract
Frequent itemsets (or frequent pattern) mining, which is the basic step during data stream mining, has been paid more and more attention by researchers. Because of the uncertainties and continuities of data stream, the time-efficiency and space-efficiency of many mining algorithms are unaccepted. In this paper, hashed table is introduced to represent the synoptic data structure. By this way, the memory footprints in Lossy Counting algorithms can be reduced. In addition, the algorithm of frequent itemsets mining based on D-Hashed Table (MFS-HT for short) is proposed to obtain the items whose frequency count exceeded a user-specified threshold in data streams. Comparing with Lossy Counting and a similar algorithm called Mining Frequent Item sets over data Streams by Matrix (MISM for short), the experiment result shows that MFS-HT is more effective both in time and space efficiency.
Keywords
cryptography; data mining; D-hash table; data stream mining; frequent itemsets mining; frequent pattern mining; lossy counting; Computer science; Computerized monitoring; Data engineering; Data mining; Data structures; Educational institutions; Frequency; Itemsets; Sampling methods; Software engineering; D-HT; MFS-HT; data stream; frequent itemset;
fLanguage
English
Publisher
ieee
Conference_Titel
Software Engineering, Artificial Intelligences, Networking and Parallel/Distributed Computing, 2009. SNPD '09. 10th ACIS International Conference on
Conference_Location
Daegu
Print_ISBN
978-0-7695-3642-2
Type
conf
DOI
10.1109/SNPD.2009.29
Filename
5286661
Link To Document