Title :
A modified cutoff scanning matrix protein representation for enhancing protein function prediction
Author :
Maghawry, Huda A. ; Mostafa, Mostafa G. M. ; Abdul-Aziz, Mohamed H. ; Gharib, Tarek F.
Author_Institution :
Fac. of Comput. & Inf. Sci., Ain Shams Univ., Cairo, Egypt
Abstract :
Protein function prediction is an active research area in bioinformatics. Protein functions are highly related to their structures. Therefore, effective structure based protein representations are required. Pires et al. [BMC Genomics, 12, S12 (2011)] proposed a cutoff scanning matrix (CSM) method for protein representation that utilizes distance patterns between protein residues and a maximum cutoff. This paper proposes a modified cutoff scanning matrix (MCSM) representation for enhancing protein function prediction. The proposed representation considers the whole protein instead of using cutoff. A comparative analysis was done to evaluate the proposed MCSM method and the original CSM method. Two different classification algorithms, Random Forest and K-nearest neighbor (KNN), were used in the analysis. The aspect of protein function considered is based on enzyme activity. The results show that the proposed MCSM representation outperforms the CSM representation with a prediction accuracy of 90.12% and 80.27% for superfamily and family level, respectively, with accuracy improvement of about 5 % on average.
Keywords :
bioinformatics; data structures; enzymes; matrix algebra; pattern classification; KNN; MCSM; bioinformatics; classification algorithm; enzyme activity; k-nearest neighbor algorithm; modified cutoff scanning matrix; protein function prediction; protein representation; random forest algorithm; Accuracy; Atomic clocks; Carbon; Computers; Educational institutions; Proteins; Vectors; cutoff scanning matrix; pattern analysis and classification; protein function prediction; protein structure representation;
Conference_Titel :
Informatics and Systems (INFOS), 2014 9th International Conference on
Conference_Location :
Cairo
Print_ISBN :
978-977-403-689-7
DOI :
10.1109/INFOS.2014.7036706