Skip to main navigation Skip to search Skip to main content

Fast-protein-cluster: Parallel and optimized clustering of large-scale protein modeling data

  • University of Washington

Research output: Contribution to journalArticlepeer-review

10 Scopus citations

Abstract

Motivation: fast-protein-cluster is a fast, parallel and memory efficient package used to cluster 60 000 sets of protein models (with up to 550 000 models per set) generated by the Nutritious Rice for the World project. Results: fast-protein-cluster is an optimized and extensible toolkit that supports Root Mean Square Deviation after optimal superposition (RMSD) and Template Modeling score (TM-score) as metrics. RMSD calculations using a laptop CPU are 60times; faster than qcprot and 3 faster than current graphics processing unit (GPU) implementations. New GPU code further increases the speed of RMSD and TM-score calculations. fast-protein-cluster provides novel k-means and hierarchical clustering methods that are up to 250×and 2000times; faster, respectively, than Clusco, and identify significantly more accurate models than Spicker and Clusco.

Original languageEnglish
Pages (from-to)1774-1776
Number of pages3
JournalBioinformatics
Volume30
Issue number12
DOIs
StatePublished - Jun 15 2014

Fingerprint

Dive into the research topics of 'Fast-protein-cluster: Parallel and optimized clustering of large-scale protein modeling data'. Together they form a unique fingerprint.

Cite this