DocumentCode :
2465007
Title :
Cluster-based KNN missing value imputation for DNA microarray data
Author :
Keerin, Phimmarin ; Kurutach, Werasak ; Boongoen, Tossapon
Author_Institution :
Fac. of Inf. Sci. & Technol., Mahanakorn Univ. of Technol., Bangkok, Thailand
fYear :
2012
fDate :
14-17 Oct. 2012
Firstpage :
445
Lastpage :
450
Abstract :
Gene expressions measured using microarrays usually encounter the problem of missing values. Leaving this unsolved may critically degrade the reliability of any consequent down-stream analysis or medical application. Yet, a further study of microarray data might be impossible with many analysis methods requiring a complete data set. This paper introduces a new methodology to impute missing values in microarray data. The proposed algorithm, CKNN impute, is an extension of k nearest neighbor imputation with local data clustering being incorporated for improved quality and efficiency. Gene expression data is typically represented as a matrix whose rows and columns correspond to genes and experiments, respectively. CKNN kicks off by finding a complete dataset via the removal of rows with missing value(s). Then, k clusters and their corresponding centroids are obtained by applying a clustering technique on the complete dataset. A set of similar genes of the target gene (with missing values) are those belonging to the cluster, whose centroid is the closest the target. Having known this, the target gene is imputed by applying k nearest neighbor method with similar genes previously determined. Empirical evaluation with published gene expression datasets suggest that the proposed technique performs better than the classical k nearest neighbor method and its extension found in the literature.
Keywords :
biology computing; data analysis; matrix algebra; molecular biophysics; pattern clustering; CKNN impute algorithm; DNA microarray data; cluster-based KNN missing value imputation; clustering technique; data analysis method; down-stream analysis; gene expression; k-nearest neighbor; matrix representation; medical application; Algorithm design and analysis; Cancer; Clustering algorithms; Euclidean distance; Gene expression; Humans; Partitioning algorithms; clustering; imputation; microarray data; missing value;
fLanguage :
English
Publisher :
ieee
Conference_Titel :
Systems, Man, and Cybernetics (SMC), 2012 IEEE International Conference on
Conference_Location :
Seoul
Print_ISBN :
978-1-4673-1713-9
Electronic_ISBN :
978-1-4673-1712-2
Type :
conf
DOI :
10.1109/ICSMC.2012.6377764
Filename :
6377764
Link To Document :
بازگشت