Skip to main navigation Skip to search Skip to main content

Canary: Decentralized Distributed Deep Learning Via Gradient Sketch and Partition in Multi-Interface Networks

  • Qihua Zhou
  • , Kun Wang
  • , Haodong Lu
  • , Wenyao Xu
  • , Yanfei Sun
  • , Song Guo
  • Nanjing University of Posts and Telecommunications
  • University of California at Los Angeles
  • Hong Kong Polytechnic University

Research output: Contribution to journalArticlepeer-review

6 Scopus citations

Abstract

The multi-interface networks are efficient infrastructures to deploy distributed Deep Learning (DL) tasks as the model gradients generated by each worker can be exchanged to others via different links in parallel. Although this decentralized parameter synchronization mechanism can reduce the time of gradient exchange, building a high-performance distributed DL architecture still requires the balance of communication efficiency and computational utilization, i.e., addressing the issues of traffic burst, data consistency, and programming convenience. To achieve this goal, we intend to asynchronously exchange gradient pieces without the central control in multi-interface networks. We propose the Piece-level Gradient Exchange and Multi-interface Collective Communication to handle parameter synchronization and traffic transmission, respectively. Specifically, we design the gradient sketch approach based on 8-bit uniform quantization to compress gradient tensors and introduce the colayer abstraction to better handle gradient partition, exchange and pipelining. Also, we provide general programming interfaces to capture the synchronization semantics and build the Gradient Exchange Index (GEI) data structures to make our approach online applicable. We implement our algorithms into a prototype system called ${\sf Canary}$Canary by using PyTorch-1.4.0. Experiments conducted in Alibaba Cloud demonstrate that ${\sf Canary}$Canary reduces 56.28 percent traffic on average and completes the training by up to 1.61x, 2.28x, and 2.84x faster than BML, Ako on PyTorch, and PS on TensorFlow, respectively.

Original languageEnglish
Article number9252115
Pages (from-to)900-917
Number of pages18
JournalIEEE Transactions on Parallel and Distributed Systems
Volume32
Issue number4
DOIs
StatePublished - Apr 1 2021

Keywords

  • decentralized architecture
  • deep learning
  • Distributed systems
  • gradient sketch
  • multi-interface network

Fingerprint

Dive into the research topics of 'Canary: Decentralized Distributed Deep Learning Via Gradient Sketch and Partition in Multi-Interface Networks'. Together they form a unique fingerprint.

Cite this