Skip to main navigation Skip to search Skip to main content

Replay attack detection based on distortion by loudspeaker for voice authentication

  • Wuhan University
  • CAS - Institute of Computing Technology

Research output: Contribution to journalArticlepeer-review

16 Scopus citations

Abstract

Identity authentication based on Automatic Speaker Verification (ASV) has attracted extensive attention. Voice can be used as a substitute of password in many applications. However, the security of current ASV systems has been seriously challenged by many malicious spoofing attacks. Among all those attacks, replay attack is one of the biggest threats to the ASV System, where an adversary can use a pre-recorded speech sample of the legal user to access the ASV system. In this paper, we present a replay attack detection (RAD) scheme to distinguish normal speech and replayed speech. We focus on the distortion caused by loudspeaker: low-frequency attenuation and high-frequency harmonics, and present a suite of RAD features DL-RAD, including Harmonic Energy Ratio (HER), Low Spectral Ratio (LSR), Low Spectral Variance (LSV), and Low Spectral Difference Variance (LSDV), to describe the different characteristics between the normal speech signal and replay speech signal. SVM is adopted as a classifier to evaluate the performance of these features. Experiment results show that the True Positive Rate (TPR), True Negative Rate (TNR) of the proposed method are about 98.15% and 98.75% respectively, which are significantly better than the existing scheme. The proposed scheme can be applied to both text-dependent and text-independent ASV systems.

Original languageEnglish
Pages (from-to)8383-8396
Number of pages14
JournalMultimedia Tools and Applications
Volume78
Issue number7
DOIs
StatePublished - Apr 1 2019

Keywords

  • Automatic Speaker Verification (ASV)
  • Loudspeaker
  • Low-frequency attenuation
  • Replay Attack Detection (RAD)
  • Spoofing attack

Fingerprint

Dive into the research topics of 'Replay attack detection based on distortion by loudspeaker for voice authentication'. Together they form a unique fingerprint.

Cite this