TY - GEN
T1 - Automatic extraction of deep phenotypes for precision medicine in chronic kidney disease
AU - Singh, Prerna
AU - Chandola, Varun
AU - Fox, Chester
N1 - Publisher Copyright:
© 2017 Association for Computing Machinery.
PY - 2017/7/2
Y1 - 2017/7/2
N2 - Chronic Kidney Disease (CKD) is one of the deadliest diseases in the world, with 10% of the global population affected by the disease. Identifying subpopulations with characteristic disease progressions is important to find more efficient treatments for patients with this disease. The abundance of electronic health records (EHR) data can be used to find meaningful subtypes for CKD but comes with challenges during analysis, including irregular data sampling, and skewness in the data collected over time. In this paper, multiple regression techniques were used to fill in the missing estimated glomerular filtration rate (or EGFR - a key measure for kidney function) trajectory data, so it can be clustered effectively. Clustering is applied to the enhanced data to obtain six subtypes, which capture crucial trends in the disease progression of patients. Moreover, the characteristics of patients in each of the subtypes had minor differences from others. These characteristics demonstrate risk factors and positive lifestyles choices of patients with CKD, which can help develop new treatments for CKD.
AB - Chronic Kidney Disease (CKD) is one of the deadliest diseases in the world, with 10% of the global population affected by the disease. Identifying subpopulations with characteristic disease progressions is important to find more efficient treatments for patients with this disease. The abundance of electronic health records (EHR) data can be used to find meaningful subtypes for CKD but comes with challenges during analysis, including irregular data sampling, and skewness in the data collected over time. In this paper, multiple regression techniques were used to fill in the missing estimated glomerular filtration rate (or EGFR - a key measure for kidney function) trajectory data, so it can be clustered effectively. Clustering is applied to the enhanced data to obtain six subtypes, which capture crucial trends in the disease progression of patients. Moreover, the characteristics of patients in each of the subtypes had minor differences from others. These characteristics demonstrate risk factors and positive lifestyles choices of patients with CKD, which can help develop new treatments for CKD.
KW - Partitioning around Medoids
KW - Regression
KW - Spline
KW - Time-series clustering
UR - https://www.scopus.com/pages/publications/85025459492
U2 - 10.1145/3079452.3079489
DO - 10.1145/3079452.3079489
M3 - Conference contribution
AN - SCOPUS:85025459492
T3 - ACM International Conference Proceeding Series
SP - 195
EP - 199
BT - DH 2017 - Proceedings of the 2017 International Conference on Digital Health
PB - Association for Computing Machinery
T2 - 7th International Conference on Digital Health, DH 2017
Y2 - 2 July 2017 through 5 July 2017
ER -