Skip to main navigation Skip to search Skip to main content

Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection

  • Cai Yu
  • , Shan Jia
  • , Xiaomeng Fu
  • , Jin Liu
  • , Jiahe Tian
  • , Jiao Dai
  • , Xi Wang
  • , Siwei Lyu
  • , Jizhong Han
  • CAS - Institute of Information Engineering
  • University of Chinese Academy of Sciences
  • SUNY Buffalo
  • CAS - Institute of Microelectronics

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

9 Scopus citations

Abstract

With the rising prevalence of deepfakes, there is a growing interest in developing generalizable detection methods for various types of deepfakes. While effective in their specific modalities, traditional detection methods fall short in addressing the generalizability of detection across diverse cross-modal deepfakes. This paper aims to explicitly learn potential cross-modal correlation to enhance deepfake detection towards various generation scenarios. Our approach introduces a correlation distillation task, which models the inherent cross-modal correlation based on content information. This strategy helps to prevent the model from overfitting merely to audio-visual synchronization. Additionally, we present the Cross-Modal Deepfake Dataset (CMDFD), a comprehensive dataset with four generation methods to evaluate the detection of diverse cross-modal deepfakes. The experimental results on CMDFD and FakeAVCeleb datasets demonstrate the superior generalizability of our method over existing state-of-the-art methods. Our code and data can be found at https://github.com/ljj898/CMDFD-Dataset-and-Deepfake-Detection.

Original languageEnglish
Title of host publication2024 IEEE International Conference on Multimedia and Expo, ICME 2024
PublisherIEEE Computer Society
ISBN (Electronic)9798350390155
DOIs
StatePublished - 2024
Event2024 IEEE International Conference on Multimedia and Expo, ICME 2024 - Niagra Falls, Canada
Duration: Jul 15 2024Jul 19 2024

Publication series

NameProceedings - IEEE International Conference on Multimedia and Expo
ISSN (Print)1945-7871
ISSN (Electronic)1945-788X

Conference

Conference2024 IEEE International Conference on Multimedia and Expo, ICME 2024
Country/TerritoryCanada
CityNiagra Falls
Period07/15/2407/19/24

Keywords

  • audio-visual deepfake detection
  • Multimedia forensics
  • multimodal learning

Fingerprint

Dive into the research topics of 'Explicit Correlation Learning for Generalizable Cross-Modal Deepfake Detection'. Together they form a unique fingerprint.

Cite this