TY - GEN
T1 - ParallelSFL
T2 - 30th International Conference on Mobile Computing and Networking, ACM MobiCom 2024
AU - Liao, Yunming
AU - Xu, Yang
AU - Xu, Hongli
AU - Yao, Zhiwei
AU - Huang, Liusheng
AU - Qiao, Chunming
N1 - Publisher Copyright:
© 2024 Copyright is held by the owner/author(s). Publication rights licensed to ACM.
PY - 2024/12/4
Y1 - 2024/12/4
N2 - Mobile devices contribute more than half of the world's web traffic, providing massive and diverse data for powering various federated learning (FL) applications. In order to avoid the communication bottleneck on the parameter server (PS) and accelerate the training of large-scale models on resource-constraint workers in edge computing (EC) system, we propose a novel split federated learning (SFL) framework, termed ParallelSFL. Concretely, we split an entire model into a bottom submodel and a top submodel, and divide participating workers into multiple clusters, each of which collaboratively performs the SFL training procedure and exchanges entire models with the PS. However, considering the statistical and system heterogeneity in edge systems, it is challenging to arrange suitable workers to specific clusters for efficient model training. To address these challenges, we carefully develop an effective clustering strategy by optimizing a utility function related to training efficiency and model accuracy. Specifically, ParallelSFL partitions workers into different clusters under the heterogeneity restrictions, thereby promoting model accuracy as well as training efficiency. Meanwhile, ParallelSFL assigns diverse and appropriate local updating frequencies for each cluster to further address system heterogeneity. Extensive experiments are conducted on a physical platform with 80 NVIDIA Jetson devices, and the experimental results show that ParallelSFL can reduce the traffic consumption by at least 21%, speed up the model training by at least 1.36X, and improve model accuracy by at least 5% in heterogeneous scenarios, compared to the baselines.
AB - Mobile devices contribute more than half of the world's web traffic, providing massive and diverse data for powering various federated learning (FL) applications. In order to avoid the communication bottleneck on the parameter server (PS) and accelerate the training of large-scale models on resource-constraint workers in edge computing (EC) system, we propose a novel split federated learning (SFL) framework, termed ParallelSFL. Concretely, we split an entire model into a bottom submodel and a top submodel, and divide participating workers into multiple clusters, each of which collaboratively performs the SFL training procedure and exchanges entire models with the PS. However, considering the statistical and system heterogeneity in edge systems, it is challenging to arrange suitable workers to specific clusters for efficient model training. To address these challenges, we carefully develop an effective clustering strategy by optimizing a utility function related to training efficiency and model accuracy. Specifically, ParallelSFL partitions workers into different clusters under the heterogeneity restrictions, thereby promoting model accuracy as well as training efficiency. Meanwhile, ParallelSFL assigns diverse and appropriate local updating frequencies for each cluster to further address system heterogeneity. Extensive experiments are conducted on a physical platform with 80 NVIDIA Jetson devices, and the experimental results show that ParallelSFL can reduce the traffic consumption by at least 21%, speed up the model training by at least 1.36X, and improve model accuracy by at least 5% in heterogeneous scenarios, compared to the baselines.
KW - edge computing
KW - split federated learning
KW - statistical heterogeneity
KW - system heterogeneity
UR - https://www.scopus.com/pages/publications/105002390941
U2 - 10.1145/3636534.3690665
DO - 10.1145/3636534.3690665
M3 - Conference contribution
AN - SCOPUS:105002390941
T3 - ACM MobiCom 2024 - Proceedings of the 30th International Conference on Mobile Computing and Networking
SP - 845
EP - 860
BT - ACM MobiCom 2024 - Proceedings of the 30th International Conference on Mobile Computing and Networking
PB - Association for Computing Machinery, Inc
Y2 - 18 November 2024 through 22 November 2024
ER -