• DocumentCode
    1878326
  • Title

    Mining Approximate Frequency Itemsets over Data Streams Based on D-Hash Table

  • Author

    Ju, Chunhua ; You, Gang

  • Author_Institution
    Coll. of Comput. Sci. & Inf. Eng., Zhejiang Gongshang Univ., Hangzhou, China
  • fYear
    2009
  • fDate
    27-29 May 2009
  • Firstpage
    249
  • Lastpage
    254
  • Abstract
    Frequent itemsets (or frequent pattern) mining, which is the basic step during data stream mining, has been paid more and more attention by researchers. Because of the uncertainties and continuities of data stream, the time-efficiency and space-efficiency of many mining algorithms are unaccepted. In this paper, hashed table is introduced to represent the synoptic data structure. By this way, the memory footprints in Lossy Counting algorithms can be reduced. In addition, the algorithm of frequent itemsets mining based on D-Hashed Table (MFS-HT for short) is proposed to obtain the items whose frequency count exceeded a user-specified threshold in data streams. Comparing with Lossy Counting and a similar algorithm called Mining Frequent Item sets over data Streams by Matrix (MISM for short), the experiment result shows that MFS-HT is more effective both in time and space efficiency.
  • Keywords
    cryptography; data mining; D-hash table; data stream mining; frequent itemsets mining; frequent pattern mining; lossy counting; Computer science; Computerized monitoring; Data engineering; Data mining; Data structures; Educational institutions; Frequency; Itemsets; Sampling methods; Software engineering; D-HT; MFS-HT; data stream; frequent itemset;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Software Engineering, Artificial Intelligences, Networking and Parallel/Distributed Computing, 2009. SNPD '09. 10th ACIS International Conference on
  • Conference_Location
    Daegu
  • Print_ISBN
    978-0-7695-3642-2
  • Type

    conf

  • DOI
    10.1109/SNPD.2009.29
  • Filename
    5286661