TY - GEN
T1 - Implementing a Gaussian process learning algorithm in mixed parallel environment
AU - Chandola, Varun
AU - Vatsavai, Ranga Raju
PY - 2011
Y1 - 2011
N2 - In this paper, we present a scalability analysis of a parallel Gaussian process training algorithm to simultaneously analyze a massive number of time series. We study three different parallel implementations: using threads, MPI, and a hybrid implementation using threads and MPI. We compare the scalability for the multi-threaded implementation on three different hardware platforms: a Mac desktop with two quad-core Intel Xeon processors (16 virtual cores), a Linux cluster node with four quad-core 2.3 GHz AMD Opteron processors, and SGI Altix ICE 8200 cluster node with two quad-core Intel Xeon processors (16 virtual cores). We also study the scalability of the MPI based and the hybrid MPI and thread based implementations on the SGI cluster with 128 nodes (2048 cores). Experimental results show that the hybrid implementation scales better than the multi-threaded and MPI based implementations. The application of the proposed algorithm is demonstrated in analyzing massive remote sensing observation data. The hybrid implementation, using 1536 cores, can analyze a data set with over 4 million time series in nearly 5 seconds while the serial algorithm takes nearly 12 hours to process the same data set.
AB - In this paper, we present a scalability analysis of a parallel Gaussian process training algorithm to simultaneously analyze a massive number of time series. We study three different parallel implementations: using threads, MPI, and a hybrid implementation using threads and MPI. We compare the scalability for the multi-threaded implementation on three different hardware platforms: a Mac desktop with two quad-core Intel Xeon processors (16 virtual cores), a Linux cluster node with four quad-core 2.3 GHz AMD Opteron processors, and SGI Altix ICE 8200 cluster node with two quad-core Intel Xeon processors (16 virtual cores). We also study the scalability of the MPI based and the hybrid MPI and thread based implementations on the SGI cluster with 128 nodes (2048 cores). Experimental results show that the hybrid implementation scales better than the multi-threaded and MPI based implementations. The application of the proposed algorithm is demonstrated in analyzing massive remote sensing observation data. The hybrid implementation, using 1536 cores, can analyze a data set with over 4 million time series in nearly 5 seconds while the serial algorithm takes nearly 12 hours to process the same data set.
KW - gaussian process
KW - scalability
KW - time series
UR - https://www.scopus.com/pages/publications/84857554241
U2 - 10.1145/2133173.2133176
DO - 10.1145/2133173.2133176
M3 - Conference contribution
AN - SCOPUS:84857554241
SN - 9781450311809
T3 - ScalA'11 - Proceedings of the 2011 ACM Workshop on Scalable Algorithms for Large-Scale Systems, Co-located with SC'11
SP - 3
EP - 6
BT - ScalA'11 - Proceedings of the 2011 ACM Workshop on Scalable Algorithms for Large-Scale Systems, Co-located with SC'11
T2 - 2011 ACM Workshop on Scalable Algorithms for Large-Scale Systems, ScalA'11, Co-located with SC'11
Y2 - 14 November 2011 through 14 November 2011
ER -