DocumentCode :
3341258
Title :
Dataset threshold for the performance estimators in supervised machine learning experiments
Author :
Omary, Z. ; Mtenzi, F.
Author_Institution :
Sch. of Comput., Dublin Inst. of Technol., Dublin, Ireland
fYear :
2009
fDate :
9-12 Nov. 2009
Firstpage :
1
Lastpage :
8
Abstract :
The establishment of dataset threshold is one among the first steps when comparing the performance of machine learning algorithms. It involves the use of different datasets with different sample sizes in relation to the number of attributes and the number of instances available in the dataset. Currently, there is no limit which has been set for those who are unfamiliar with machine learning experiments on the categorisation of these datasets, as either small or large, based on the two factors. In this paper we perform experiments in order to establish dataset threshold. The established dataset threshold will help unfamiliar supervised machine learning experimenters to categorize datasets based on the number of instances and attributes and then choose the appropriate performance estimation method. The experiments will involve the use of four different datasets from UCI machine learning repository and two performance estimators. The performance of the methods will be measured using f1-score.
Keywords :
learning (artificial intelligence); dataset threshold; performance estimation method; supervised machine learning; Communications technology; Context; Decision trees; Error analysis; Knowledge acquisition; Machine learning; Machine learning algorithms; Performance evaluation; Radio frequency; Support vector machines;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Internet Technology and Secured Transactions, 2009. ICITST 2009. International Conference for
Conference_Location :
London
Print_ISBN :
978-1-4244-5647-5
Type :
conf
DOI :
10.1109/ICITST.2009.5402500
Filename :
5402500
Link To Document :
بازگشت