Abstract
Motivation: fast-protein-cluster is a fast, parallel and memory efficient package used to cluster 60 000 sets of protein models (with up to 550 000 models per set) generated by the Nutritious Rice for the World project. Results: fast-protein-cluster is an optimized and extensible toolkit that supports Root Mean Square Deviation after optimal superposition (RMSD) and Template Modeling score (TM-score) as metrics. RMSD calculations using a laptop CPU are 60times; faster than qcprot and 3 faster than current graphics processing unit (GPU) implementations. New GPU code further increases the speed of RMSD and TM-score calculations. fast-protein-cluster provides novel k-means and hierarchical clustering methods that are up to 250×and 2000times; faster, respectively, than Clusco, and identify significantly more accurate models than Spicker and Clusco.
| Original language | English |
|---|---|
| Pages (from-to) | 1774-1776 |
| Number of pages | 3 |
| Journal | Bioinformatics |
| Volume | 30 |
| Issue number | 12 |
| DOIs | |
| State | Published - Jun 15 2014 |
Fingerprint
Dive into the research topics of 'Fast-protein-cluster: Parallel and optimized clustering of large-scale protein modeling data'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver