Title :
Lattice Histograms: a Resilient Synopsis Structure
Author :
Karras, Panagiotis ; Mamoulis, Nikos
Author_Institution :
Dept. of Inf., Univ. of Zurich, Zurich
Abstract :
Despite the surge of interest in data reduction techniques over the past years, no method has been proposed to date that can always achieve approximation quality preferable to that of the optimal plain histogram for a target error metric. In this paper, we introduce the lattice histogram: a novel data reduction method that discovers and exploits any arbitrary hierarchy in the data, and achieves approximation quality provably at least as high as an optimal histogram for any data reduction problem. We formulate LH construction techniques with approximation guarantees for general error metrics. We show that the case of minimizing a maximum-error metric can be solved by a specialized, memory-sparing approach; we exploit this solution to design reduced-space heuristics for the general- error case. We develop a mixed synopsis approach, applicable to the space-efficient high-quality summarization of very large data sets. We experimentally corroborate the superiority of LHs in approximation quality over previous techniques with representative error metrics and diverse data sets.
Keywords :
data reduction; very large databases; arbitrary data hierarchy; data reduction; lattice histograms; maximum-error metric; memory-sparing approach; resilient synopsis structure; very large data sets; Compaction; Computer science; Data mining; Decision support systems; Histograms; Indexing; Informatics; Lattices; Quality management; Query processing;
Conference_Titel :
Data Engineering, 2008. ICDE 2008. IEEE 24th International Conference on
Conference_Location :
Cancun
Print_ISBN :
978-1-4244-1836-7
Electronic_ISBN :
978-1-4244-1837-4
DOI :
10.1109/ICDE.2008.4497433