• DocumentCode
    2403955
  • Title

    Data cleaning and XML: the DBLP experience

  • Author

    Low, Wai Lup ; Hyong, T. Wee ; Lee, Mong Li ; Ling, Tok Wang

  • Author_Institution
    Sch. of Comput., Nat. Univ. of Singapore, Singapore
  • fYear
    2002
  • fDate
    2002
  • Firstpage
    269
  • Abstract
    With the increasing popularity of data-centric XML, data warehousing and mining applications are being developed for rapidly burgeoning XML data repositories. Data quality will no doubt be a critical factor for the success of such applications. Data cleaning, which refers to the processes used to improve data quality, has been well researched in the context of traditional databases. In earlier work we developed a knowledge-based framework for data cleaning relational databases. In this work, we present a novel attempt to apply this framework to XML databases. Our experimental dataset is the DBLP database, a popular online XML bibliography database used by many researchers
  • Keywords
    data integrity; data mining; data warehouses; hypermedia markup languages; DBLP database; XML data repositories; data cleaning; data mining; data quality; data warehousing; data-centric XML; knowledge-based framework; online XML bibliography database; relational databases; Cleaning; Data engineering; XML;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Engineering, 2002. Proceedings. 18th International Conference on
  • Conference_Location
    San Jose, CA
  • ISSN
    1063-6382
  • Print_ISBN
    0-7695-1531-2
  • Type

    conf

  • DOI
    10.1109/ICDE.2002.994723
  • Filename
    994723