Abstract
In this paper we discuss the challenge of equitably combining continuous (quantita-tive) and categorical (qualitative) variables for the purpose of cluster analysis. Existing techniques require strong parametric assumptions, or difficult-to-specify tuning parameters. We describe the kamila package, which includes a weighted k-means approach to clustering mixed-type data, a method for estimating weights for mixed-type data (Modha-Spangler weighting), and an additional semiparametric method recently proposed in the literature (KAMILA). We include a discussion of strategies for estimating the number of clusters in the data, and describe the implementation of one such method in the current R package. Background and usage of these clustering methods are presented. We then show how the KAMILA algorithm can be adapted to a map-reduce framework, and implement the resulting algorithm using Hadoop for clustering very large mixed-type data sets.
| Original language | English |
|---|---|
| Journal | Journal of Statistical Software |
| Volume | 83 |
| DOIs | |
| State | Published - 2018 |
Keywords
- Clustering
- Hadoop
- Kernel density estimator
- Mixed data
- Mixed-type data
- Mixture model
- R
- Semiparametric
- Unsupervised learning
Fingerprint
Dive into the research topics of 'kamila: Clustering mixed-type data in R and hadoop'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver