TY - GEN
T1 - Private Data Imputation
AU - Kati, Addelkarim
AU - Kerschbaum, Florian
AU - Blanton, Marina
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - Data imputation is an important data preparation task where the data analyst replaces missing or erroneous values to increase the expected accuracy of downstream analyses. The accuracy improvement of data imputation extends to private data analyses across distributed databases. However, existing data imputation methods violate the privacy of the data rendering the privacy protection in the downstream analyses obsolete. We conclude that private data analysis requires private data imputation. In this paper, we present the first optimized protocols for private data imputation. We consider the case of horizontally and vertically split data sets. Our optimization aims to reduce most of the computation to private set intersection (or at least oblivious programmable pseudo-random function) protocols which can be very efficiently computed. We show that private data imputation has - on average across all evaluated datasets - an accuracy advantage of 20% in case of vertically split data and 5% in case of horizontally split data over imputing data locally. In case of the worst data split we observed that imputing using our method resulted in an accuracy improvement (Root Mean Square Error reduction) of up to 32.7 times over the vertically split data and 3.4 times in case of horizontally split data. Our protocols are very efficient and run in 2.4 seconds in case of vertically split data and 8.4 seconds in case of horizontally split data for 100,000 records evaluated in the 10 Gbps network setting, performing one data imputation.
AB - Data imputation is an important data preparation task where the data analyst replaces missing or erroneous values to increase the expected accuracy of downstream analyses. The accuracy improvement of data imputation extends to private data analyses across distributed databases. However, existing data imputation methods violate the privacy of the data rendering the privacy protection in the downstream analyses obsolete. We conclude that private data analysis requires private data imputation. In this paper, we present the first optimized protocols for private data imputation. We consider the case of horizontally and vertically split data sets. Our optimization aims to reduce most of the computation to private set intersection (or at least oblivious programmable pseudo-random function) protocols which can be very efficiently computed. We show that private data imputation has - on average across all evaluated datasets - an accuracy advantage of 20% in case of vertically split data and 5% in case of horizontally split data over imputing data locally. In case of the worst data split we observed that imputing using our method resulted in an accuracy improvement (Root Mean Square Error reduction) of up to 32.7 times over the vertically split data and 3.4 times in case of horizontally split data. Our protocols are very efficient and run in 2.4 seconds in case of vertically split data and 8.4 seconds in case of horizontally split data for 100,000 records evaluated in the 10 Gbps network setting, performing one data imputation.
UR - https://www.scopus.com/pages/publications/105044405502
U2 - 10.1109/SP63933.2026.00127
DO - 10.1109/SP63933.2026.00127
M3 - Conference contribution
AN - SCOPUS:105044405502
T3 - Proceedings - IEEE Symposium on Security and Privacy
SP - 2057
EP - 2075
BT - Proceedings - 47th IEEE Symposium on Security and Privacy, SP 2026
A2 - Oprea, Alina
A2 - Nita-Rotaru, Cristina
A2 - Papernot, Nicolas
PB - Institute of Electrical and Electronics Engineers Inc.
T2 - 47th IEEE Symposium on Security and Privacy, SP 2026
Y2 - 18 May 2026 through 21 May 2026
ER -