TY - GEN
T1 - Modality-Agnostic Deepfakes Detection
AU - Cai, Yu
AU - Chen, Peng
AU - Tian, Jiahe
AU - Liu, Jin
AU - Dai, Jiao
AU - Wang, Xi
AU - Jia, Shan
AU - Lyu, Siwei
AU - Han, Jizhong
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s). Publication rights licensed to ACM.
PY - 2025/6/17
Y1 - 2025/6/17
N2 - As AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated. While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked. Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities. However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable. In such cases, audio-visual detection methods are less practical than two independent unimodal methods. Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector. In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms. To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR) as a preliminary task. This efficiently extracts speech correlations across modalities, a feature challenging for deepfakes to replicate. Additionally, we propose a dual-label detection approach that follows the structure of AVSR to support the independent detection of each modality. Extensive experiments on three audio-visual datasets show that our scheme outperforms state-of-the-art detection methods with promising performance on modality-agnostic audio/video deepfakes.
AB - As AI-generated content (AIGC) thrives, deepfakes have expanded from single-modality falsification to cross-modal fake content creation, where either audio or visual components can be manipulated. While using two unimodal detectors can detect audio-visual deepfakes, cross-modal forgery clues could be overlooked. Existing multimodal deepfake detectors typically establish correspondence between the audio and visual modalities for binary real/fake classification and require the co-occurrence of both modalities. However, in real-world multi-modal applications, missing modality scenarios may occur where either modality is unavailable. In such cases, audio-visual detection methods are less practical than two independent unimodal methods. Consequently, the detector can not always obtain the number or type of manipulated modalities beforehand, necessitating a fake-modality-agnostic audio-visual detector. In this work, we introduce a comprehensive framework that is agnostic to fake modalities, which facilitates the identification of multimodal deepfakes and handles situations with missing modalities, regardless of the manipulations embedded in audio, video, or even cross-modal forms. To enhance the modeling of cross-modal forgery clues, we employ audio-visual speech recognition (AVSR) as a preliminary task. This efficiently extracts speech correlations across modalities, a feature challenging for deepfakes to replicate. Additionally, we propose a dual-label detection approach that follows the structure of AVSR to support the independent detection of each modality. Extensive experiments on three audio-visual datasets show that our scheme outperforms state-of-the-art detection methods with promising performance on modality-agnostic audio/video deepfakes.
KW - Forgery Detection.
KW - Multimedia Forensics
KW - Multimodal Learning
UR - https://www.scopus.com/pages/publications/105031778748
U2 - 10.1145/3733102.3733133
DO - 10.1145/3733102.3733133
M3 - Conference contribution
AN - SCOPUS:105031778748
T3 - IHandMMSec 2025 - Proceedings of the 2025 ACM Workshop on Information Hiding and Multimedia Security
SP - 12
EP - 23
BT - IHandMMSec 2025 - Proceedings of the 2025 ACM Workshop on Information Hiding and Multimedia Security
A2 - Agarwal, Shruti
A2 - Craver, Scott
A2 - Jia, Shan
A2 - Wong, Chau-Wai
A2 - Tondi, Benedetta
PB - Association for Computing Machinery, Inc
T2 - 13th ACM Workshop on Information Hiding and Multimedia Security, IHandMMSec 2025
Y2 - 18 June 2025 through 20 June 2025
ER -