DocumentCode
173339
Title
Classification of a real live heart failure clinical dataset- Is TAN Bayes better than other Bayes?
Author
Moore, L. ; Kambhampati, C. ; Cleland, J.G.F.
Author_Institution
Dept. of Comput. Sci., Univ. of Hull, Kingston upon Hull, UK
fYear
2014
fDate
5-8 Oct. 2014
Firstpage
882
Lastpage
887
Abstract
Real live clinical data often present itself with a number of usual challenges, such as class imbalance, high dimensionality and missing data. There is the added complexity of the data being distributed non-uniformly and skewed. Thus the performance of classical classification methods with this type of data is lower than with other types of data. Classification based on Bayes is often suggested as a better method, however, the typical assumption made for Bayes, such as variable and data distributions are not satisfied by real clinical data. This paper focuses on improving the performance of Bayesian classifiers but also on how the underlying structures of the data affects the performance. Thus this paper will focus on Bayesian methodologies, namely use of non-parametric Kernel Density Estimation (KDE) and Tree Augmented Naïve Bayes (TAN). The aim is to measure the performance on the heart failure dataset and by focusing on how the data structure improves the classification. The missing data present in the clinical heart failure datasets are replaced using two imputation methods and results compared. We also apply the imputed datasets on three classifiers including J48 (decision tree), naïve Bayesian multinomial and Bayesian network. The experiments show an improvement on the naïve Bayes using KDE, however TAN achieves significant improvement with the different missing value imputation methods. It is seen that TAN not only improves performance of the classifier, but also enhances prediction accuracy while maintaining efficiency and model simplicity.
Keywords
Bayes methods; bioinformatics; data structures; decision trees; learning (artificial intelligence); pattern classification; Bayesian classifiers; Bayesian methodologies; Bayesian network; J48; KDE; TAN Bayes; classification methods; data distributions; data structure; heart failure clinical dataset; kernel density estimation; naive Bayesian multinomial; tree augmented naive Bayes; Bayes methods; Complexity theory; Estimation; Gaussian distribution; Heart; Kernel; Support vector machines; classification; distribution; heart failure; kernel density estimation; naive Bayes; tree augmented naïve Bayes;
fLanguage
English
Publisher
ieee
Conference_Titel
Systems, Man and Cybernetics (SMC), 2014 IEEE International Conference on
Conference_Location
San Diego, CA
Type
conf
DOI
10.1109/SMC.2014.6974023
Filename
6974023
Link To Document