• DocumentCode
    3292132
  • Title

    Efficient Clustering-Based Outlier Detection Algorithm for Dynamic Data Stream

  • Author

    Elahi, Manzoor ; Li, Kun ; Nisar, Wasif ; Lv, Xinjie ; Wang, Hongan

  • Author_Institution
    Intell. Eng. Lab., Inst. of Software Chinese Acad. of Sci., Beijing
  • Volume
    5
  • fYear
    2008
  • fDate
    18-20 Oct. 2008
  • Firstpage
    298
  • Lastpage
    304
  • Abstract
    Anomaly detection is currently an important and active research problem in many fields and involved in numerous applications. Most of the existing methods are based on distance measure. But in case of data stream these methods are not very efficient as computational point of view. Most of the exiting work on outlier detection in data stream declare a point as an outlier/inlier as soon as it arrive due to limited memory resources as compared to the huge data stream, to declare an outlier as it arrive often can lead us to a wrong decision, because of dynamic nature of the incoming data. In this paper we introduced a clustering based approach, which divide the stream in chunks and cluster each chunk using k-mean in fixed number of clusters. Instead of keeping only the summary information, which often used in case of clustering data stream, we keep the candidate outliers and mean value of every cluster for the next fixed number of steam chunks, to make sure that the detected candidate outliers are the real outliers. By employing the mean value of the clusters of previous chunk with mean values of the current chunk of stream, we decide better outlierness for data stream objects. Several experiments on different dataset confirm that our technique can find better outliers with low computational cost than the other exiting distance based approaches of outlier detection in data stream.
  • Keywords
    data mining; security of data; anomaly detection; clustering based approach; clustering-based outlier detection algorithm; dynamic data stream; Application software; Computational efficiency; Data engineering; Data mining; Databases; Detection algorithms; Fuzzy systems; Information technology; Knowledge engineering; Nearest neighbor searches; Clustering; Data Stream; Outlier detection;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Fuzzy Systems and Knowledge Discovery, 2008. FSKD '08. Fifth International Conference on
  • Conference_Location
    Jinan Shandong
  • Print_ISBN
    978-0-7695-3305-6
  • Type

    conf

  • DOI
    10.1109/FSKD.2008.374
  • Filename
    4666541