TY - GEN
T1 - DHC
T2 - 3rd IEEE Symposium on BioInformatics and BioEngineering, BIBE 2003
AU - Jiang, Daxin
AU - Pei, Jian
AU - Zhang, Aidong
N1 - Publisher Copyright:
© 2003 IEEE.
PY - 2003
Y1 - 2003
N2 - Clustering the time series gene expression data is an important task in bioinformatics research and biomedical applications. Recently, some clustering methods have been adapted or proposed. However, some concerns still remain, such as the robustness of the mining methods, as well as the quality and the interpretability of the mining results. In this paper, we tackle the problem of effectively clustering time series gene expression data by proposing algorithm DHC, a density-based, hierarchical clustering method. We use a density-based approach to identify the clusters such that the clustering results are of high quality and robustness. Moreover, The mining result is in the form of a density tree, which uncovers the embedded clusters in a data set. The inner-structures, the borders and the outliers of the clusters can be further investigated using the attraction tree, which is an intermediate result of the mining. By these two trees, the internal structure of the data set can be visualized effectively. Our empirical evaluation using some real-world data sets show that the method is effective, robust and scalable. It matches the ground truth provided by bioinformatics experts very well in the sample data sets.
AB - Clustering the time series gene expression data is an important task in bioinformatics research and biomedical applications. Recently, some clustering methods have been adapted or proposed. However, some concerns still remain, such as the robustness of the mining methods, as well as the quality and the interpretability of the mining results. In this paper, we tackle the problem of effectively clustering time series gene expression data by proposing algorithm DHC, a density-based, hierarchical clustering method. We use a density-based approach to identify the clusters such that the clustering results are of high quality and robustness. Moreover, The mining result is in the form of a density tree, which uncovers the embedded clusters in a data set. The inner-structures, the borders and the outliers of the clusters can be further investigated using the attraction tree, which is an intermediate result of the mining. By these two trees, the internal structure of the data set can be visualized effectively. Our empirical evaluation using some real-world data sets show that the method is effective, robust and scalable. It matches the ground truth provided by bioinformatics experts very well in the sample data sets.
UR - https://www.scopus.com/pages/publications/84942597099
U2 - 10.1109/BIBE.2003.1188978
DO - 10.1109/BIBE.2003.1188978
M3 - Conference contribution
AN - SCOPUS:84942597099
T3 - Proceedings - 3rd IEEE Symposium on BioInformatics and BioEngineering, BIBE 2003
SP - 393
EP - 400
BT - Proceedings - 3rd IEEE Symposium on BioInformatics and BioEngineering, BIBE 2003
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 10 March 2003 through 12 March 2003
ER -