TY - GEN
T1 - mmCLIP
T2 - 22nd ACM Conference on Embedded Networked Sensor Systems, SenSys 2024
AU - Cao, Qiming
AU - Xue, Hongfei
AU - Liu, Tianci
AU - Wang, Xingchen
AU - Wang, Haoyu
AU - Zhang, Xincheng
AU - Su, Lu
N1 - Publisher Copyright:
© 2024 Copyright is held by the owner/author(s).
PY - 2024/11/4
Y1 - 2024/11/4
N2 - Millimeter-wave (mmWave) based human activity recognition (HAR) systems have demonstrated promising performance in various applications, leveraging the power of deep neural networks. However, these systems are suffering from the scarcity of available mmWave data for model training. To address this challenge, we explore the possibility of transferring knowledge from large AI models built on massive text and visual data to enhance the generalizability of mmWave-based HAR models. Towards this end, we introduce mmCLIP, a novel system that aligns mmWave signal space and text space to facilitate zero-shot recognition for unseen activities. To enable this alignment, we employ cross-modality signal synthesis to augment mmWave signal data using large human mesh datasets and design an activity attribute decomposition and recomposition approach to characterize the semantic interconnections among activities. We conducted extensive experiments to demonstrate the effectiveness of our proposed framework.
AB - Millimeter-wave (mmWave) based human activity recognition (HAR) systems have demonstrated promising performance in various applications, leveraging the power of deep neural networks. However, these systems are suffering from the scarcity of available mmWave data for model training. To address this challenge, we explore the possibility of transferring knowledge from large AI models built on massive text and visual data to enhance the generalizability of mmWave-based HAR models. Towards this end, we introduce mmCLIP, a novel system that aligns mmWave signal space and text space to facilitate zero-shot recognition for unseen activities. To enable this alignment, we employ cross-modality signal synthesis to augment mmWave signal data using large human mesh datasets and design an activity attribute decomposition and recomposition approach to characterize the semantic interconnections among activities. We conducted extensive experiments to demonstrate the effectiveness of our proposed framework.
KW - human activity recognition
KW - large language model
KW - mmwave
KW - signal augmentation
KW - visual-language model
KW - wireless sensing
UR - https://www.scopus.com/pages/publications/85211789626
U2 - 10.1145/3666025.3699331
DO - 10.1145/3666025.3699331
M3 - Conference contribution
AN - SCOPUS:85211789626
T3 - SenSys 2024 - Proceedings of the 2024 ACM Conference on Embedded Networked Sensor Systems
SP - 184
EP - 197
BT - SenSys 2024 - Proceedings of the 2024 ACM Conference on Embedded Networked Sensor Systems
PB - Association for Computing Machinery, Inc
Y2 - 4 November 2024 through 7 November 2024
ER -