An ensemble approach for large-scale identification of proteinprotein interactions using the alignments of multiple sequences

Lei Wang, Zhu Hong You, Xing Chen, Jian Qiang Li, Xin Yan, Wei Zhang, Yu An Huang

Research output: Contribution to journalArticlepeer-review

40 Scopus citations

Abstract

Protein-Protein Interactions (PPI) is not only the critical component of various biological processes in cells, but also the key to understand the mechanisms leading to healthy and diseased states in organisms. However, it is time-consuming and costintensive to identify the interactions among proteins using biological experiments. Hence, how to develop a more efficient computational method rapidly became an attractive topic in the post-genomic era. In this paper, we propose a novel method for inference of protein-protein interactions from protein amino acids sequences only. Specifically, protein amino acids sequence is firstly transformed into Position- Specific Scoring Matrix (PSSM) generated by multiple sequences alignments; then the Pseudo PSSM is used to extract feature descriptors. Finally, ensemble Rotation Forest (RF) learning system is trained to predict and recognize PPIs based solely on protein sequence feature. When performed the proposed method on the three benchmark data sets (Yeast, H. pylori, and independent dataset) for predicting PPIs, our method can achieve good average accuracies of 98.38%, 89.75%, and 96.25%, respectively. In order to further evaluate the prediction performance, we also compare the proposed method with other methods using same benchmark data sets. The experiment results demonstrate that the proposed method consistently outperforms other state-of-the-art method. Therefore, our method is effective and robust and can be taken as a useful tool in exploring and discovering new relationships between proteins. A web server is made publicly available at the URL http://202.119.201.126:8888/PsePSSM/ for academic use.

Original languageEnglish
Pages (from-to)5149-5159
Number of pages11
JournalOncotarget
Volume8
Issue number3
DOIs
StatePublished - 2017
Externally publishedYes

Keywords

  • Cancer
  • Disease
  • Multiple sequences alignments
  • Position-specific scoring matrix

Fingerprint

Dive into the research topics of 'An ensemble approach for large-scale identification of proteinprotein interactions using the alignments of multiple sequences'. Together they form a unique fingerprint.

Cite this