Skip to main navigation Skip to search Skip to main content

Wavoice: A Noise-resistant Multi-modal Speech Recognition System Fusing mmWave and Audio Signals

  • Tiantian Liu
  • , Ming Gao
  • , Feng Lin
  • , Chao Wang
  • , Zhongjie Ba
  • , Jinsong Han
  • , Wenyao Xu
  • , Kui Ren
  • Zhejiang University
  • Key Laboratory of Blockchain and Cyberspace Governance of Zhejiang Province

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

87 Scopus citations

Abstract

With the advance in automatic speech recognition, voice user interface has gained popularity recently. Since the COVID-19 pandemic, VUI is increasingly preferred in online communication due to its non-contact. Additionally, various ambient noise impedes the public applications of voice user interfaces due to the requirement of audio-only speech recognition methods for a high signal-to-noise ratio. In this paper, we present Wavoice, the first noise-resistant multi-modal speech recognition system that fuses two distinct voice sensing modalities, i.e., millimeter-wave (mmWave) signals and audio signals from a microphone, together. One key contribution is that we model the inherent correlation between mmWave and audio signals. Based on it, Wavoice facilitates the real-time noise-resistant voice activity detection and user targeting from multiple speakers. Furthermore, we elaborate on two novel modules into the neural attention mechanism for multi-modal signals fusion, and result in accurate speech recognition. Extensive experiments verify Wavoice's effectiveness under various conditions with the character recognition error rate below 1% in a range of 7 meters. Wavoice outperforms existing audio-only speech recognition methods with lower character error rate and word error rate. The evaluation in complex scenes validates the robustness of Wavoice.

Original languageEnglish
Title of host publicationSenSys 2021 - Proceedings of the 2021 19th ACM Conference on Embedded Networked Sensor Systems
PublisherAssociation for Computing Machinery, Inc
Pages97-110
Number of pages14
ISBN (Electronic)9781450390972
DOIs
StatePublished - Nov 15 2021
Event19th ACM Conference on Embedded Networked Sensor Systems, SenSys 2021 - Coimbra, Portugal
Duration: Nov 15 2021Nov 17 2021

Publication series

NameSenSys 2021 - Proceedings of the 2021 19th ACM Conference on Embedded Networked Sensor Systems

Conference

Conference19th ACM Conference on Embedded Networked Sensor Systems, SenSys 2021
Country/TerritoryPortugal
CityCoimbra
Period11/15/2111/17/21

Keywords

  • mmWave sensing
  • multimodal fusion
  • Speech recognition
  • voice user interface

Fingerprint

Dive into the research topics of 'Wavoice: A Noise-resistant Multi-modal Speech Recognition System Fusing mmWave and Audio Signals'. Together they form a unique fingerprint.

Cite this