Title :
Cluster tree based multi-label classification for protein function prediction
Author :
Qingyao Wu ; Yunming Ye ; Xiaofeng Zhang ; Shen-Shyang Ho
Author_Institution :
Dept. of Comput. Sci., Harbin Inst. of Technol., Shenzhen, China
Abstract :
Automatically assigning functions for unknown proteins is a key task in computational biology. Proteins in nature have multiple classes according to the functions they perform. Many efforts have been made to cast the protein function prediction into a multi-label learning problem. This paper proposes a novel Cluster Tree based Multi-label Learning algorithm (CTML) for protein function prediction. The main idea is to compute a set of predictive labels associated at each node for multi-label prediction by using the k-means clustering techniques and the predictive functions via the learning data at the nodes. With the propagation of the predictive labels from the root node to the leaf node, the correlations between labels can be preserved. Experimental results on benchmark data (genbase and yeast datasets) show that the proposed CTML algorithm is effective in predicting protein functions. Moreover, the classification performance of the CTML algorithm is competitive against the other baseline multi-label learning algorithms.
Keywords :
bioinformatics; learning (artificial intelligence); pattern clustering; proteins; CTML method; Cluster Tree based Multi-label Learning algorithm; computational biology; genbase dataset; k-means clustering technique; protein function prediction; yeast dataset; Clustering algorithms; Measurement; Partitioning algorithms; Prediction algorithms; Protein engineering; Proteins; Vectors; Data mining; Multi-label classification; Multi-label data; Protein function prediction;
Conference_Titel :
Bioinformatics and Biomedicine (BIBM), 2013 IEEE International Conference on
Conference_Location :
Shanghai
DOI :
10.1109/BIBM.2013.6732548