TY - GEN
T1 - Wavoice
T2 - 19th ACM Conference on Embedded Networked Sensor Systems, SenSys 2021
AU - Liu, Tiantian
AU - Gao, Ming
AU - Lin, Feng
AU - Wang, Chao
AU - Ba, Zhongjie
AU - Han, Jinsong
AU - Xu, Wenyao
AU - Ren, Kui
N1 - Publisher Copyright:
© 2021 ACM.
PY - 2021/11/15
Y1 - 2021/11/15
N2 - With the advance in automatic speech recognition, voice user interface has gained popularity recently. Since the COVID-19 pandemic, VUI is increasingly preferred in online communication due to its non-contact. Additionally, various ambient noise impedes the public applications of voice user interfaces due to the requirement of audio-only speech recognition methods for a high signal-to-noise ratio. In this paper, we present Wavoice, the first noise-resistant multi-modal speech recognition system that fuses two distinct voice sensing modalities, i.e., millimeter-wave (mmWave) signals and audio signals from a microphone, together. One key contribution is that we model the inherent correlation between mmWave and audio signals. Based on it, Wavoice facilitates the real-time noise-resistant voice activity detection and user targeting from multiple speakers. Furthermore, we elaborate on two novel modules into the neural attention mechanism for multi-modal signals fusion, and result in accurate speech recognition. Extensive experiments verify Wavoice's effectiveness under various conditions with the character recognition error rate below 1% in a range of 7 meters. Wavoice outperforms existing audio-only speech recognition methods with lower character error rate and word error rate. The evaluation in complex scenes validates the robustness of Wavoice.
AB - With the advance in automatic speech recognition, voice user interface has gained popularity recently. Since the COVID-19 pandemic, VUI is increasingly preferred in online communication due to its non-contact. Additionally, various ambient noise impedes the public applications of voice user interfaces due to the requirement of audio-only speech recognition methods for a high signal-to-noise ratio. In this paper, we present Wavoice, the first noise-resistant multi-modal speech recognition system that fuses two distinct voice sensing modalities, i.e., millimeter-wave (mmWave) signals and audio signals from a microphone, together. One key contribution is that we model the inherent correlation between mmWave and audio signals. Based on it, Wavoice facilitates the real-time noise-resistant voice activity detection and user targeting from multiple speakers. Furthermore, we elaborate on two novel modules into the neural attention mechanism for multi-modal signals fusion, and result in accurate speech recognition. Extensive experiments verify Wavoice's effectiveness under various conditions with the character recognition error rate below 1% in a range of 7 meters. Wavoice outperforms existing audio-only speech recognition methods with lower character error rate and word error rate. The evaluation in complex scenes validates the robustness of Wavoice.
KW - mmWave sensing
KW - multimodal fusion
KW - Speech recognition
KW - voice user interface
UR - https://www.scopus.com/pages/publications/85120892228
U2 - 10.1145/3485730.3485945
DO - 10.1145/3485730.3485945
M3 - Conference contribution
AN - SCOPUS:85120892228
T3 - SenSys 2021 - Proceedings of the 2021 19th ACM Conference on Embedded Networked Sensor Systems
SP - 97
EP - 110
BT - SenSys 2021 - Proceedings of the 2021 19th ACM Conference on Embedded Networked Sensor Systems
PB - Association for Computing Machinery, Inc
Y2 - 15 November 2021 through 17 November 2021
ER -