DocumentCode
2403955
Title
Data cleaning and XML: the DBLP experience
Author
Low, Wai Lup ; Hyong, T. Wee ; Lee, Mong Li ; Ling, Tok Wang
Author_Institution
Sch. of Comput., Nat. Univ. of Singapore, Singapore
fYear
2002
fDate
2002
Firstpage
269
Abstract
With the increasing popularity of data-centric XML, data warehousing and mining applications are being developed for rapidly burgeoning XML data repositories. Data quality will no doubt be a critical factor for the success of such applications. Data cleaning, which refers to the processes used to improve data quality, has been well researched in the context of traditional databases. In earlier work we developed a knowledge-based framework for data cleaning relational databases. In this work, we present a novel attempt to apply this framework to XML databases. Our experimental dataset is the DBLP database, a popular online XML bibliography database used by many researchers
Keywords
data integrity; data mining; data warehouses; hypermedia markup languages; DBLP database; XML data repositories; data cleaning; data mining; data quality; data warehousing; data-centric XML; knowledge-based framework; online XML bibliography database; relational databases; Cleaning; Data engineering; XML;
fLanguage
English
Publisher
ieee
Conference_Titel
Data Engineering, 2002. Proceedings. 18th International Conference on
Conference_Location
San Jose, CA
ISSN
1063-6382
Print_ISBN
0-7695-1531-2
Type
conf
DOI
10.1109/ICDE.2002.994723
Filename
994723
Link To Document