Skip to main navigation Skip to search Skip to main content

LAUGHING HYENA DISTILLERY Extracting Compact Recurrences From Convolutions

  • Stefano Massaroli
  • , Michael Poli
  • , Daniel Y. Fu
  • , Hermann Kumbong
  • , Rom N. Parnichkun
  • , Aman Timalsina
  • , David W. Romero
  • , Quinn McIntyre
  • , Beidi Chen
  • , Atri Rudra
  • , Ce Zhang
  • , Christopher Re
  • , Stefano Ermon
  • , Yoshua Bengio
  • Université de Montréal
  • Stanford University
  • The University of Tokyo
  • Purdue University
  • Vrije Universiteit Amsterdam
  • Carnegie Mellon University
  • The University of Chicago

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

15 Scopus citations

Abstract

Recent advances in attention-free sequence models rely on convolutions as alternatives to the attention operator at the core of Transformers. In particular, long convolution sequence models have achieved state-of-the-art performance in many domains, but incur a significant cost during auto-regressive inference workloads - naively requiring a full pass (or caching of activations) over the input sequence for each generated token - similarly to attention-based models. In this paper, we seek to enable O(1) compute and memory cost per token in any pre-trained long convolution architecture to reduce memory footprint and increase throughput during generation. Concretely, our methods consist in extracting low-dimensional linear state-space models from each convolution layer, building upon rational interpolation and model-order reduction techniques. We further introduce architectural improvements to convolution-based layers such as Hyena: by weight-tying the filters across channels into heads, we achieve higher pretraining quality and reduce the number of filters to be distilled. The resulting model achieves 10× higher throughput than Transformers and 1.5× higher than Hyena at 1.3B parameters, without any loss in quality after distillation.

Original languageEnglish
Title of host publicationAdvances in Neural Information Processing Systems 36 - 37th Conference on Neural Information Processing Systems, NeurIPS 2023
EditorsA. Oh, T. Neumann, A. Globerson, K. Saenko, M. Hardt, S. Levine
PublisherNeural information processing systems foundation
ISBN (Electronic)9781713899921
StatePublished - 2023
Event37th Conference on Neural Information Processing Systems, NeurIPS 2023 - New Orleans, United States
Duration: Dec 10 2023Dec 16 2023

Publication series

NameAdvances in Neural Information Processing Systems
Volume36
ISSN (Print)1049-5258

Conference

Conference37th Conference on Neural Information Processing Systems, NeurIPS 2023
Country/TerritoryUnited States
CityNew Orleans
Period12/10/2312/16/23

Fingerprint

Dive into the research topics of 'LAUGHING HYENA DISTILLERY Extracting Compact Recurrences From Convolutions'. Together they form a unique fingerprint.

Cite this