Skip to main navigation Skip to search Skip to main content

What Do Audio Transformers Hear? Probing Their Representations For Language Delivery & Structure

  • SUNY Buffalo
  • Indraprastha Institute of Information Technology Delhi

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

32 Scopus citations

Abstract

Transformer models across multiple domains such as natural language processing and speech form an unavoidable part of the tech stack of practitioners and researchers alike. Au-dio transformers that exploit representational learning to train on unlabeled speech have recently been used for tasks from speaker verification to discourse-coherence with much success. However, little is known about what these models learn and represent in the high-dimensional latent space. In this paper, we interpret two such recent state-of-the-art models, wav2vec2.0 and Mockingjay, on linguistic and acoustic features. We probe each of their layers to understand what it is learning and at the same time, we draw a distinction between the two models. By comparing their performance across a wide variety of settings including native, non-native, read and spontaneous speeches, we also show how much these models are able to learn transferable features. Our results show that the models are capable of significantly capturing a wide range of characteristics such as audio, fluency, supraseg-mental pronunciation, and even syntactic and semantic text-based characteristics. For each category of characteristics, we identify a learning pattern for each framework and conclude which model and which layer of that model is better for a specific category of feature to choose for feature extraction for downstream tasks.

Original languageEnglish
Title of host publicationProceedings - 22nd IEEE International Conference on Data Mining Workshops, ICDMW 2022
EditorsK. Selcuk Candan, Thang N. Dinh, My T. Thai, Takashi Washio
PublisherIEEE Computer Society
Pages910-925
Number of pages16
ISBN (Electronic)9798350346091
DOIs
StatePublished - 2022
Event22nd IEEE International Conference on Data Mining Workshops, ICDMW 2022 - Orlando, United States
Duration: Nov 28 2022Dec 1 2022

Publication series

NameIEEE International Conference on Data Mining Workshops, ICDMW
Volume2022-November
ISSN (Print)2375-9232
ISSN (Electronic)2375-9259

Conference

Conference22nd IEEE International Conference on Data Mining Workshops, ICDMW 2022
Country/TerritoryUnited States
CityOrlando
Period11/28/2212/1/22

Keywords

  • Audio Transformers
  • Interpretability
  • Language Delivery
  • Language Structure
  • Transformers
  • wav2vec2.0

Fingerprint

Dive into the research topics of 'What Do Audio Transformers Hear? Probing Their Representations For Language Delivery & Structure'. Together they form a unique fingerprint.

Cite this