DocumentCode
2457620
Title
Random Error Reduction in Similarity Search on Time Series: A Statistical Approach
Author
Wu, Wush Chi-Hsuan ; Yeh, Mi-Yen ; Pei, Jian
Author_Institution
Inst. of Inf. Sci., Acad. Sinica, Taipei, Taiwan
fYear
2012
fDate
1-5 April 2012
Firstpage
858
Lastpage
869
Abstract
Errors in measurement can be categorized into two types: systematic errors that are predictable, and random errors that are inherently unpredictable and have null expected value. Random error is always present in a measurement. More often than not, readings in time series may contain inherent random errors due to causes like dynamic error, drift, noise, hysteresis, digitalization error and limited sampling frequency. Random errors may affect the quality of time series analysis substantially. Unfortunately, most of the existing time series mining and analysis methods, such as similarity search, clustering, and classification tasks, do not address random errors, possibly because random error in a time series, which can be modeled as a random variable of unknown distribution, is hard to handle. In this paper, we tackle this challenging problem. Taking similarity search as an example, which is an essential task in time series analysis, we develop MISQ, a statistical approach for random error reduction in time series analysis. The major intuition in our method is to use only the readings at different time instants in a time series to reduce random errors. We achieve a highly desirable property in MISQ: it can ensure that the recall is above a user-specified threshold. An extensive empirical study on 20 benchmark real data sets clearly shows that our method can lead to better performance than the baseline method without random error reduction in real applications such as classification. Moreover, MISQ achieves good quality in similarity search.
Keywords
data mining; pattern classification; pattern clustering; time series; MISQ; classification task; clustering task; digitalization error; drift; dynamic error; hysteresis; limited sampling frequency; noise; random error reduction; similarity search; statistical approach; systematic error; time series analysis; time series mining; user-specified threshold; Error analysis; Measurement uncertainty; Random variables; Reactive power; Time measurement; Time series analysis;
fLanguage
English
Publisher
ieee
Conference_Titel
Data Engineering (ICDE), 2012 IEEE 28th International Conference on
Conference_Location
Washington, DC
ISSN
1063-6382
Print_ISBN
978-1-4673-0042-1
Type
conf
DOI
10.1109/ICDE.2012.83
Filename
6228139
Link To Document