DocumentCode
2416602
Title
Empirical Evaluation of Cost Overrun Prediction with Imbalance Data
Author
Tsunoda, Masateru ; Monden, Akito ; Shibata, Jun-ichiro ; Matsumoto, Ken-ichi
fYear
2011
fDate
16-18 May 2011
Firstpage
415
Lastpage
420
Abstract
To prevent cost overrun of software projects, it is necessary for project managers to identify projects which have high risk of cost overrun in the early phase. So far, discriminant methods such as linear discriminant analysis and logistic regression have been used to predict cost overrun projects. However, accuracy of discriminant methods often becomes low when a dataset used for predict is imbalanced, i.e. there exists a large difference between the number of cost overrun projects and non cost overrun projects. In this paper, we compared accuracy of linear discriminant analysis, logistic regression, classification tree, Mahalanobis-Taguchi method, and collaborative filtering, by changing the percentage of cost overrun projects in the dataset. The result showed that collaborative filtering was highest accuracy among five methods. When the number of cost overrun projects and non cost overrun was balanced in the dataset, linear discriminant analysis was second highest accuracy, and when it was not balanced, Mahalanobis-Taguchi method was second highest among five methods.
Keywords
Accuracy; Classification tree analysis; Collaboration; Filtering; Logistics; Mathematical model; Predictive models; Collaborative Filtering; Mahalanobis-Taguchi method; biased data; failure prone project; risk management;
fLanguage
English
Publisher
ieee
Conference_Titel
Computer and Information Science (ICIS), 2011 IEEE/ACIS 10th International Conference on
Conference_Location
Sanya, China
Print_ISBN
978-1-4577-0141-2
Type
conf
DOI
10.1109/ICIS.2011.71
Filename
6086505
Link To Document