An Efficient Approach for Outlier Detection with Imperfect Data Labels

Author

Bo Liu ; Yanshan Xiao ; Yu, Philip S. ; Zhifeng Hao ; Longbing Cao

Author_Institution

Dept. of Autom., Guangdong Univ. of Technol., Guangzhou, China

Volume

26

Issue

7

fYear

2014

fDate

Jul-14

Firstpage

1602

Lastpage

1616

Abstract

The task of outlier detection is to identify data objects that are markedly different from or inconsistent with the normal set of data. Most existing solutions typically build a model using the normal data and identify outliers that do not fit the represented model very well. However, in addition to normal data, there also exist limited negative examples or outliers in many applications, and data may be corrupted such that the outlier detection data is imperfectly labeled. These make outlier detection far more difficult than the traditional ones. This paper presents a novel outlier detection approach to address data with imperfect labels and incorporate limited abnormal examples into learning. To deal with data with imperfect labels, we introduce likelihood values for each input data which denote the degree of membership of an example toward the normal and abnormal classes respectively. Our proposed approach works in two steps. In the first step, we generate a pseudo training dataset by computing likelihood values of each example based on its local behavior. We present kernel (k) -means clustering method and kernel LOF-based method to compute the likelihood values. In the second step, we incorporate the generated likelihood values and limited abnormal examples into SVDD-based learning framework to build a more accurate classifier for global outlier detection. By integrating local and global outlier detection, our proposed method explicitly handles data with imperfect labels and enhances the performance of outlier detection. Extensive experiments on real life datasets have demonstrated that our proposed approaches can achieve a better tradeoff between detection rate and false alarm rate as compared to state-of-the-art outlier detection approaches.

Keywords

learning (artificial intelligence); pattern classification; pattern clustering; SVDD-based learning framework; abnormal class; classifier; data object identification; global outlier detection; imperfect data labels; kernel LOF-based method; kernel-means clustering method; likelihood values; local outlier detection; membership degree; normal class; normal data; pseudo training dataset; Computational modeling; Data models; Educational institutions; Kernel; Support vector machines; Training; Training data; Outlier detection; data mining; data of uncertainty; outlier detection;

fLanguage

English

Journal_Title

Knowledge and Data Engineering, IEEE Transactions on

Publisher

ieee

ISSN

1041-4347

Type

jour

DOI

10.1109/TKDE.2013.108

Filename

6547621