TY - GEN
T1 - FLUDE
T2 - 2026 IEEE Conference on Computer Communications, INFOCOM 2026
AU - Wang, Shilong
AU - Liu, Jianchun
AU - Xu, Hongli
AU - Qiao, Chunming
N1 - Publisher Copyright:
© 2026 IEEE.
PY - 2026
Y1 - 2026
N2 - In a federated learning (FL) system, many devices, such as smartphones, are often undependable (e.g., frequently disconnected from WiFi) during training. Existing FL frameworks always assume a dependable environment and simply exclude undependable devices from training, leading to poor model performance and resource wastage. In this paper, we propose FLUDE to effectively deal with undependable environments. First, FLUDE assesses the dependability of devices based on the probability distribution of their historical behaviors (e.g., the likelihood of successfully completing training). Based on this assessment, FLUDE adaptively selects devices with high dependability for training. To mitigate resource wastage during the training phase, FLUDE maintains a model cache on each device, aiming to preserve the latest training state for later use in case local training on an undependable device is interrupted. Moreover, FLUDE proposes a staleness-aware strategy to judiciously distribute the global model to a subset of devices, thus significantly reducing resource wastage while maintaining model performance. We have implemented FLUDE on two physical platforms with 120 smartphones and NVIDIA Jetson devices. Extensive experiments show that FLUDE effectively improves model performance and resource efficiency of FL in undependable environments.
AB - In a federated learning (FL) system, many devices, such as smartphones, are often undependable (e.g., frequently disconnected from WiFi) during training. Existing FL frameworks always assume a dependable environment and simply exclude undependable devices from training, leading to poor model performance and resource wastage. In this paper, we propose FLUDE to effectively deal with undependable environments. First, FLUDE assesses the dependability of devices based on the probability distribution of their historical behaviors (e.g., the likelihood of successfully completing training). Based on this assessment, FLUDE adaptively selects devices with high dependability for training. To mitigate resource wastage during the training phase, FLUDE maintains a model cache on each device, aiming to preserve the latest training state for later use in case local training on an undependable device is interrupted. Moreover, FLUDE proposes a staleness-aware strategy to judiciously distribute the global model to a subset of devices, thus significantly reducing resource wastage while maintaining model performance. We have implemented FLUDE on two physical platforms with 120 smartphones and NVIDIA Jetson devices. Extensive experiments show that FLUDE effectively improves model performance and resource efficiency of FL in undependable environments.
KW - Device Undependability
KW - Federated Learning
KW - Machine Learning Systems
UR - https://www.scopus.com/pages/publications/105044542446
U2 - 10.1109/INFOCOM59046.2026.11571430
DO - 10.1109/INFOCOM59046.2026.11571430
M3 - Conference contribution
AN - SCOPUS:105044542446
T3 - Proceedings - IEEE INFOCOM
BT - INFOCOM 2026 - IEEE Conference on Computer Communications
PB - Institute of Electrical and Electronics Engineers Inc.
Y2 - 18 May 2026 through 21 May 2026
ER -