DocumentCode
3189383
Title
Predicting and Optimizing Classifier Utility with the Power Law
Author
Last, Mark
fYear
2007
fDate
28-31 Oct. 2007
Firstpage
219
Lastpage
224
Abstract
When data collection is costly and/or takes a significant amount of time, an early prediction of the classifier performance is extremely important for the design of the data mining process. Power law has been shown in the past to be a good predictor of decision- tree error rates as a function of the sample size. In this paper, we show that the optimal training set size for a given dataset can be computed from a learning curve characterized by a power law. Such a curve can be approximated using a small subset of potentially available data and then used to estimate the expected trade-off between the error rate and the amount of additional observations. The proposed approach to projected optimization of classifier utility is demonstrated and evaluated on several benchmark datasets.
Keywords
Accuracy; Classification algorithms; Conferences; Cost function; Data mining; Design optimization; Error analysis; Predictive models; Sampling methods; Training data;
fLanguage
English
Publisher
ieee
Conference_Titel
Data Mining Workshops, 2007. ICDM Workshops 2007. Seventh IEEE International Conference on
Conference_Location
Omaha, NE
Print_ISBN
978-0-7695-3019-2
Electronic_ISBN
978-0-7695-3033-8
Type
conf
DOI
10.1109/ICDMW.2007.31
Filename
4476671
Link To Document