• DocumentCode
    2081410
  • Title

    PIP: A database system for great and small expectations

  • Author

    Kennedy, Oliver ; Koch, Christoph

  • Author_Institution
    Dept. of Comput. Sci., Cornell Univ., Ithaca, NY, USA
  • fYear
    2010
  • fDate
    1-6 March 2010
  • Firstpage
    157
  • Lastpage
    168
  • Abstract
    Estimation via sampling out of highly selective join queries is well known to be problematic, most notably in online aggregation. Without goal-directed sampling strategies, samples falling outside of the selection constraints lower estimation efficiency at best, and cause inaccurate estimates at worst This problem appears in general probabilistic database systems, where query processing is tightly coupled with sampling. By committing to a set of samples before evaluating the query, the engine wastes effort on samples that will be discarded, query processing that may need to be repeated, or unnecessarily large numbers of samples. We describe PIP, a general probabilistic database system that uses symbolic representations of probabilistic data to defer computation of expectations, moments, and other statistical measures until the expression to be measured is fully known. This approach is sufficiently general to admit both continuous and discrete distributions. Moreover, deferring sampling enables a broad range of goal-oriented sampling-based (as well as exact) integration techniques for computing expectations, allows the selection of the integration strategy most appropriate to the expression being measured, and can reduce the amount of sampling work required. We demonstrate the effectiveness of this approach by showing that even straightforward algorithms can make use of the added information. These algorithms have a profoundly positive impact on the efficiency and accuracy of expectation computations, particularly in the case of highly selective join queries.
  • Keywords
    database management systems; estimation theory; probability; query processing; sampling methods; subroutines; PIP; computing expectations; continuous distributions; discrete distributions; estimation efficiency; goal directed sampling strategies; goal-oriented sampling-based integration techniques; great expectations; online aggregation; probabilistic database systems; query processing; small expectations; Computer applications; Data mining; Database systems; Encoding; Engines; Predictive models; Probability distribution; Query processing; Sampling methods;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Engineering (ICDE), 2010 IEEE 26th International Conference on
  • Conference_Location
    Long Beach, CA
  • Print_ISBN
    978-1-4244-5445-7
  • Electronic_ISBN
    978-1-4244-5444-0
  • Type

    conf

  • DOI
    10.1109/ICDE.2010.5447879
  • Filename
    5447879