TY - GEN
T1 - ADMM-based weight pruning for real-time deep learning acceleration on mobile devices
AU - Li, Hongjia
AU - Liu, Ning
AU - Ma, Xiaolong
AU - Lin, Sheng
AU - Ye, Shaokai
AU - Zhang, Tianyun
AU - Lin, Xue
AU - Xu, Wenyao
AU - Wang, Yanzhi
N1 - Publisher Copyright:
© 2019 ACM.
PY - 2019/5/13
Y1 - 2019/5/13
N2 - Deep learning solutions are being increasingly deployed in mobile applications, at least for the inference phase. Due to the large model size and computational requirements, model compression for deep neural networks (DNNs) becomes necessary, especially considering the real-time requirement in embedded systems. In this paper, we extend the prior work on systematic DNN weight pruning using ADMM (Alternating Direction Method of Multipliers). We integrate ADMM regularization with masked mapping/retraining, thereby guaranteeing solution feasibility and providing high solution quality. Besides superior performance on representative DNN benchmarks (e.g., AlexNet, ResNet), we focus on two new applications facial emotion detection and eye tracking, and develop a top-down framework of DNN training, model compression, and acceleration in mobile devices. Experimental results show that with negligible accuracy degradation, the proposed method can achieve significant storage/memory reduction and speedup in mobile devices.
AB - Deep learning solutions are being increasingly deployed in mobile applications, at least for the inference phase. Due to the large model size and computational requirements, model compression for deep neural networks (DNNs) becomes necessary, especially considering the real-time requirement in embedded systems. In this paper, we extend the prior work on systematic DNN weight pruning using ADMM (Alternating Direction Method of Multipliers). We integrate ADMM regularization with masked mapping/retraining, thereby guaranteeing solution feasibility and providing high solution quality. Besides superior performance on representative DNN benchmarks (e.g., AlexNet, ResNet), we focus on two new applications facial emotion detection and eye tracking, and develop a top-down framework of DNN training, model compression, and acceleration in mobile devices. Experimental results show that with negligible accuracy degradation, the proposed method can achieve significant storage/memory reduction and speedup in mobile devices.
KW - Acceleration
KW - Mobile devices
KW - Neural networks
KW - Real-time
UR - https://www.scopus.com/pages/publications/85072671584
U2 - 10.1145/3299874.3319492
DO - 10.1145/3299874.3319492
M3 - Conference contribution
AN - SCOPUS:85072671584
T3 - Proceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI
SP - 501
EP - 506
BT - GLSVLSI 2019 - Proceedings of the 2019 Great Lakes Symposium on VLSI
PB - Association for Computing Machinery
T2 - 29th Great Lakes Symposium on VLSI, GLSVLSI 2019
Y2 - 9 May 2019 through 11 May 2019
ER -