• DocumentCode
    3189383
  • Title

    Predicting and Optimizing Classifier Utility with the Power Law

  • Author

    Last, Mark

  • fYear
    2007
  • fDate
    28-31 Oct. 2007
  • Firstpage
    219
  • Lastpage
    224
  • Abstract
    When data collection is costly and/or takes a significant amount of time, an early prediction of the classifier performance is extremely important for the design of the data mining process. Power law has been shown in the past to be a good predictor of decision- tree error rates as a function of the sample size. In this paper, we show that the optimal training set size for a given dataset can be computed from a learning curve characterized by a power law. Such a curve can be approximated using a small subset of potentially available data and then used to estimate the expected trade-off between the error rate and the amount of additional observations. The proposed approach to projected optimization of classifier utility is demonstrated and evaluated on several benchmark datasets.
  • Keywords
    Accuracy; Classification algorithms; Conferences; Cost function; Data mining; Design optimization; Error analysis; Predictive models; Sampling methods; Training data;
  • fLanguage
    English
  • Publisher
    ieee
  • Conference_Titel
    Data Mining Workshops, 2007. ICDM Workshops 2007. Seventh IEEE International Conference on
  • Conference_Location
    Omaha, NE
  • Print_ISBN
    978-0-7695-3019-2
  • Electronic_ISBN
    978-0-7695-3033-8
  • Type

    conf

  • DOI
    10.1109/ICDMW.2007.31
  • Filename
    4476671