DocumentCode
1381093
Title
A procedure for analyzing unbalanced datasets
Author
Kitchenham, Barbara
Author_Institution
Dept. of Comput. Sci., Keele Univ., UK
Volume
24
Issue
4
fYear
1998
fDate
4/1/1998 12:00:00 AM
Firstpage
278
Lastpage
301
Abstract
This paper describes a procedure for analyzing unbalanced datasets that include many nominal- and ordinal-scale factors. Such datasets are often found in company datasets used for benchmarking and productivity assessment. The two major problems caused by lack of balance are that the impact of factors can be concealed and that spurious impacts can be observed. These effects are examined with the help of two small artificial datasets. The paper proposes a method of forward pass residual analysis to analyze such datasets. The analysis procedure is demonstrated on the artificial datasets and then applied to the COCOMO dataset. The paper ends with a discussion of the advantages and limitations of the analysis procedure
Keywords
software metrics; COCOMO dataset; analysis of variance; benchmarking; forward pass residual analysis; productivity assessment; residual analysis; software metrics; statistical analysis; unbalanced datasets; Analysis of variance; Assembly; Cause effect analysis; Computer Society; Computer industry; Data analysis; Productivity; Software measurement; Software quality; Statistical analysis;
fLanguage
English
Journal_Title
Software Engineering, IEEE Transactions on
Publisher
ieee
ISSN
0098-5589
Type
jour
DOI
10.1109/32.677185
Filename
677185
Link To Document