Skip to main navigation Skip to search Skip to main content

ADMM-based weight pruning for real-time deep learning acceleration on mobile devices

  • Hongjia Li
  • , Ning Liu
  • , Xiaolong Ma
  • , Sheng Lin
  • , Shaokai Ye
  • , Tianyun Zhang
  • , Xue Lin
  • , Wenyao Xu
  • , Yanzhi Wang
  • Northeastern University
  • Syracuse University

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

24 Scopus citations

Abstract

Deep learning solutions are being increasingly deployed in mobile applications, at least for the inference phase. Due to the large model size and computational requirements, model compression for deep neural networks (DNNs) becomes necessary, especially considering the real-time requirement in embedded systems. In this paper, we extend the prior work on systematic DNN weight pruning using ADMM (Alternating Direction Method of Multipliers). We integrate ADMM regularization with masked mapping/retraining, thereby guaranteeing solution feasibility and providing high solution quality. Besides superior performance on representative DNN benchmarks (e.g., AlexNet, ResNet), we focus on two new applications facial emotion detection and eye tracking, and develop a top-down framework of DNN training, model compression, and acceleration in mobile devices. Experimental results show that with negligible accuracy degradation, the proposed method can achieve significant storage/memory reduction and speedup in mobile devices.

Original languageEnglish
Title of host publicationGLSVLSI 2019 - Proceedings of the 2019 Great Lakes Symposium on VLSI
PublisherAssociation for Computing Machinery
Pages501-506
Number of pages6
ISBN (Electronic)9781450362528
DOIs
StatePublished - May 13 2019
Event29th Great Lakes Symposium on VLSI, GLSVLSI 2019 - Tysons Corner, United States
Duration: May 9 2019May 11 2019

Publication series

NameProceedings of the ACM Great Lakes Symposium on VLSI, GLSVLSI

Conference

Conference29th Great Lakes Symposium on VLSI, GLSVLSI 2019
Country/TerritoryUnited States
CityTysons Corner
Period05/9/1905/11/19

Keywords

  • Acceleration
  • Mobile devices
  • Neural networks
  • Real-time

Fingerprint

Dive into the research topics of 'ADMM-based weight pruning for real-time deep learning acceleration on mobile devices'. Together they form a unique fingerprint.

Cite this