Skip to main navigation Skip to search Skip to main content

Parallel framework for dimensionality reduction of large-scale datasets

  • Sai Kiranmayee Samudrala
  • , Jaroslaw Zola
  • , Srinivas Aluru
  • , Baskar Ganapathysubramanian
  • Georgia Institute of Technology
  • Iowa State University

Research output: Contribution to journalArticlepeer-review

12 Scopus citations

Abstract

Dimensionality reduction refers to a set of mathematical techniques used to reduce complexity of the original high-dimensional data, while preserving its selected properties. Improvements in simulation strategies and experimental data collection methods are resulting in a deluge of heterogeneous and high-dimensional data, which often makes dimensionality reduction the only viable way to gain qualitative and quantitative understanding of the data. However, existing dimensionality reduction software often does not scale to datasets arising in real-life applications, which may consist of thousands of points with millions of dimensions. In this paper, we propose a parallel framework for dimensionality reduction of large-scale data. We identify key components underlying the spectral dimensionality reduction techniques, and propose their efficient parallel implementation. We show that the resulting framework can be used to process datasets consisting of millions of points when executed on a 16,000-core cluster, which is beyond the reach of currently availablemethods. To further demonstrate applicability of our framework we performdimensionality reduction of 75,000 images representing morphology evolution duringmanufacturing of organic solar cells in order to identify how processing parameters affect morphology evolution.

Original languageEnglish
Article number180214
JournalScientific Programming
Volume2015
DOIs
StatePublished - 2015

Fingerprint

Dive into the research topics of 'Parallel framework for dimensionality reduction of large-scale datasets'. Together they form a unique fingerprint.

Cite this