Skip to main navigation Skip to search Skip to main content

A semiparametric method for clustering mixed data

  • SUNY Buffalo
  • IBM

Research output: Contribution to journalArticlepeer-review

77 Scopus citations

Abstract

Despite the existence of a large number of clustering algorithms, clustering remains a challenging problem. As large datasets become increasingly common in a number of different domains, it is often the case that clustering algorithms must be applied to heterogeneous sets of variables, creating an acute need for robust and scalable clustering methods for mixed continuous and categorical scale data. We show that current clustering methods for mixed-type data are generally unable to equitably balance the contribution of continuous and categorical variables without strong parametric assumptions. We develop KAMILA (KAy-means for MIxed LArge data), a clustering method that addresses this fundamental problem directly. We study theoretical aspects of our method and demonstrate its effectiveness in a series of Monte Carlo simulation studies and a set of real-world applications.

Original languageEnglish
Pages (from-to)419-458
Number of pages40
JournalMachine Learning
Volume105
Issue number3
DOIs
StatePublished - Dec 1 2016

Keywords

  • Big data
  • Clustering
  • Finite mixture models
  • Mixed data
  • Unsupervised learning
  • k-means

Fingerprint

Dive into the research topics of 'A semiparametric method for clustering mixed data'. Together they form a unique fingerprint.

Cite this