• DocumentCode
    660924
  • Title

    Large-Scale RDF Dataset Slicing

  • Author

    Marx, Edgard ; Shekarpour, Saeedeh ; Auer, Stefan ; Ngomo, Axel-Cyrille Ngonga

  • Author_Institution
    Comput. Sci., Univ. of Leipzig, Leipzig, Germany
  • fYear
    2013
  • fDate
    16-18 Sept. 2013
  • Firstpage
    228
  • Lastpage
    235
  • Abstract
    In the last years an increasing number of structured data was published on the Web as Linked Open Data (LOD). Despite recent advances, consuming and using Linked Open Data within an organization is still a substantial challenge. Many of the LOD datasets are quite large and despite progress in RDF data management their loading and querying within a triple store is extremely time-consuming and resource-demanding. To overcome this consumption obstacle, we propose a process inspired by the classical Extract-Transform-Load (ETL) paradigm. In this article, we focus particularly on the selection and extraction steps of this process. We devise a fragment of SPARQL dubbed SliceSPARQL, which enables the selection of well-defined slices of datasets fulfilling typical information needs. SliceSPARQL supports graph patterns for which each connected sub graph pattern involves a maximum of one variable or IRI in its join conditions. This restriction guarantees the efficient processing of the query against a sequential dataset dump stream. As a result our evaluation shows that dataset slices can be generated an order of magnitude faster than by using the conventional approach of loading the whole dataset into a triple store and retrieving the slice by executing the query against the triple store´s SPARQL endpoint.
  • Keywords
    Internet; graph theory; organisational aspects; program slicing; query languages; ETL paradigm; LOD datasets; RDF data management; SPARQL endpoint; World Wide Web; dataset slices; extract-transform-load paradigm; graph patterns; large-scale RDF dataset slicing; linked open data; organization; sequential dataset dump stream; sliceSPARQL; structured data; subgraph pattern; Data mining; Loading; Organizations; Pattern matching; Resource description framework; Time complexity;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Semantic Computing (ICSC), 2013 IEEE Seventh International Conference on
  • Conference_Location
    Irvine, CA
  • Type

    conf

  • DOI
    10.1109/ICSC.2013.47
  • Filename
    6693522