Skip to main navigation Skip to search Skip to main content

kamila: Clustering mixed-type data in R and hadoop

  • SUNY Buffalo

Research output: Contribution to journalArticlepeer-review

69 Scopus citations

Abstract

In this paper we discuss the challenge of equitably combining continuous (quantita-tive) and categorical (qualitative) variables for the purpose of cluster analysis. Existing techniques require strong parametric assumptions, or difficult-to-specify tuning parameters. We describe the kamila package, which includes a weighted k-means approach to clustering mixed-type data, a method for estimating weights for mixed-type data (Modha-Spangler weighting), and an additional semiparametric method recently proposed in the literature (KAMILA). We include a discussion of strategies for estimating the number of clusters in the data, and describe the implementation of one such method in the current R package. Background and usage of these clustering methods are presented. We then show how the KAMILA algorithm can be adapted to a map-reduce framework, and implement the resulting algorithm using Hadoop for clustering very large mixed-type data sets.

Original languageEnglish
JournalJournal of Statistical Software
Volume83
DOIs
StatePublished - 2018

Keywords

  • Clustering
  • Hadoop
  • Kernel density estimator
  • Mixed data
  • Mixed-type data
  • Mixture model
  • R
  • Semiparametric
  • Unsupervised learning

Fingerprint

Dive into the research topics of 'kamila: Clustering mixed-type data in R and hadoop'. Together they form a unique fingerprint.

Cite this