• DocumentCode
    1260024
  • Title

    Representative Distance: A New Similarity Measure for Class Discovery From Gene Expression Data

  • Author

    Zhiwen Yu ; You, J. ; Le Li ; Hau-San Wong ; Guoqiang Han

  • Author_Institution
    Sch. of Comput. Sci. & Eng., South China Univ. of Technol., Guangzhou, China
  • Volume
    11
  • Issue
    4
  • fYear
    2012
  • Firstpage
    341
  • Lastpage
    351
  • Abstract
    Similarity measurement is one of the most important stages in the process of cancer discovery from gene expression data. Traditional distance functions, such as the Euclidean distance, the correlation coefficient measure, the cosine distance, and so on, are selected to quantify the similarity between two cancer samples. However, these measures do not take into account the properties of cancer samples and do not consider the relationships among the genes in gene expression data. In order to explore the properties of cancer samples and the relationships among genes, we design a new similarity measure called representative distance (RD) to identify cancer samples in gene expression data. Specifically, RD does not compute the distance between two cancer samples using all the genes, but only calculates the similarity using representative genes selected by the affinity propagation algorithm. Then, a similarity matrix is constructed based on the representative distance. Finally, the spectral clustering algorithm is adopted to partition the similarity matrix, and discover the biological meaningful samples. To our knowledge, this is the first time in which the representative distance is applied to class discovery for gene expression data. Experiments on real cancer datasets indicate that our similarity measure can (i) outperform most of the traditional distance measures, (ii) identify cancer samples correctly in most of the datasets.
  • Keywords
    cancer; genetics; medical computing; statistical analysis; affinity propagation algorithm; biological meaningful samples; cancer datasets; cancer discovery; cancer samples; correlation coefficient measurement; cosine distance; euclidean distance; gene expression data; representative distance; representative genes; similarity matrix; spectral clustering algorithm; traditional distance functions; traditional distance measurement; Cancer; Clustering algorithms; Correlation; Euclidean distance; Gene expression; Cancer discovery; distance; microarray; similarity measure; Algorithms; Cluster Analysis; Gene Expression Profiling; Gene Expression Regulation, Neoplastic; Neoplasms; Oligonucleotide Array Sequence Analysis;
  • fLanguage
    English
  • Journal_Title
    NanoBioscience, IEEE Transactions on
  • Publisher
    ieee
  • ISSN
    1536-1241
  • Type

    jour

  • DOI
    10.1109/TNB.2012.2208198
  • Filename
    6261551