Title :
Data cleaning and XML: the DBLP experience
Author :
Low, Wai Lup ; Hyong, T. Wee ; Lee, Mong Li ; Ling, Tok Wang
Author_Institution :
Sch. of Comput., Nat. Univ. of Singapore, Singapore
Abstract :
With the increasing popularity of data-centric XML, data warehousing and mining applications are being developed for rapidly burgeoning XML data repositories. Data quality will no doubt be a critical factor for the success of such applications. Data cleaning, which refers to the processes used to improve data quality, has been well researched in the context of traditional databases. In earlier work we developed a knowledge-based framework for data cleaning relational databases. In this work, we present a novel attempt to apply this framework to XML databases. Our experimental dataset is the DBLP database, a popular online XML bibliography database used by many researchers
Keywords :
data integrity; data mining; data warehouses; hypermedia markup languages; DBLP database; XML data repositories; data cleaning; data mining; data quality; data warehousing; data-centric XML; knowledge-based framework; online XML bibliography database; relational databases; Cleaning; Data engineering; XML;
Conference_Titel :
Data Engineering, 2002. Proceedings. 18th International Conference on
Conference_Location :
San Jose, CA
Print_ISBN :
0-7695-1531-2
DOI :
10.1109/ICDE.2002.994723