• DocumentCode
    1133296
  • Title

    Integrating Data Warehouses with Web Data: A Survey

  • Author

    Pérez, Juan Manuel ; Berlanga, Rafael ; Aramburu, María José ; Pedersen, Torben Bach

  • Author_Institution
    Dept. de Lenguajes y Sist. Informaticos, Univ. Jaume I, Castellon de la Plana
  • Volume
    20
  • Issue
    7
  • fYear
    2008
  • fDate
    7/1/2008 12:00:00 AM
  • Firstpage
    940
  • Lastpage
    955
  • Abstract
    This paper surveys the most relevant research on combining Data Warehouse (DW) and Web data. It studies the XML technologies that are currently being used to integrate, store, query and retrieve web data, and their application to DWs. The paper reviews different DW distributed architectures and the use of XML languages as an integration tool in these systems. It also introduces the problem of dealing with semi-structured data in a DW. It studies Web data repositories, the design of multidimensional databases for XML data sources and the XML extensions of On-Line Analytical Processing techniques. The paper addresses the application of information retrieval technology in a DW to exploit text-rich documents collections. The authors hope that the paper will help to discover the main limitations and opportunities that offer the combination of the DW and the Web fields, as well as, to identify open research lines.
  • Keywords
    Internet; XML; data mining; data warehouses; information retrieval; Web data repositories; XML data sources; XML extensions; XML languages; data warehouses; information retrieval technology; integration tool; multidimensional databases; online analytical processing techniques; text-rich document collections; Data warehouse and repository; XML/XSL/RDF;
  • fLanguage
    English
  • Journal_Title
    Knowledge and Data Engineering, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1041-4347
  • Type

    jour

  • DOI
    10.1109/TKDE.2007.190746
  • Filename
    4490177