DocumentCode
1260024
Title
Representative Distance: A New Similarity Measure for Class Discovery From Gene Expression Data
Author
Zhiwen Yu ; You, J. ; Le Li ; Hau-San Wong ; Guoqiang Han
Author_Institution
Sch. of Comput. Sci. & Eng., South China Univ. of Technol., Guangzhou, China
Volume
11
Issue
4
fYear
2012
Firstpage
341
Lastpage
351
Abstract
Similarity measurement is one of the most important stages in the process of cancer discovery from gene expression data. Traditional distance functions, such as the Euclidean distance, the correlation coefficient measure, the cosine distance, and so on, are selected to quantify the similarity between two cancer samples. However, these measures do not take into account the properties of cancer samples and do not consider the relationships among the genes in gene expression data. In order to explore the properties of cancer samples and the relationships among genes, we design a new similarity measure called representative distance (RD) to identify cancer samples in gene expression data. Specifically, RD does not compute the distance between two cancer samples using all the genes, but only calculates the similarity using representative genes selected by the affinity propagation algorithm. Then, a similarity matrix is constructed based on the representative distance. Finally, the spectral clustering algorithm is adopted to partition the similarity matrix, and discover the biological meaningful samples. To our knowledge, this is the first time in which the representative distance is applied to class discovery for gene expression data. Experiments on real cancer datasets indicate that our similarity measure can (i) outperform most of the traditional distance measures, (ii) identify cancer samples correctly in most of the datasets.
Keywords
cancer; genetics; medical computing; statistical analysis; affinity propagation algorithm; biological meaningful samples; cancer datasets; cancer discovery; cancer samples; correlation coefficient measurement; cosine distance; euclidean distance; gene expression data; representative distance; representative genes; similarity matrix; spectral clustering algorithm; traditional distance functions; traditional distance measurement; Cancer; Clustering algorithms; Correlation; Euclidean distance; Gene expression; Cancer discovery; distance; microarray; similarity measure; Algorithms; Cluster Analysis; Gene Expression Profiling; Gene Expression Regulation, Neoplastic; Neoplasms; Oligonucleotide Array Sequence Analysis;
fLanguage
English
Journal_Title
NanoBioscience, IEEE Transactions on
Publisher
ieee
ISSN
1536-1241
Type
jour
DOI
10.1109/TNB.2012.2208198
Filename
6261551
Link To Document