Physics-Aware Deepfake Detection via Speech-Distance Consistency

Kyeongrae Kim1, Kim Sung-Bin2, Oh Hyun-Bin2, Tae-Hyun Oh1,$\dagger$
1KAIST, 2POSTECH
Interspeech 2026
$\dagger$Corresponding Author
Teaser Image

Recent progress in generative models has made talking-head deepfakes highly realistic and easy to create, jeopardizing public trust. However, most audio-visual deepfake detectors rely on lip–speech synchronization and are designed for static, frontal videos, limiting their reliability in dynamic, in-the-wild recordings where the speaker moves and lip cues are degraded or missing.

Abstract

Recent progress in generative models has made talking-head deepfakes highly realistic and easy to create, jeopardizing public trust. However, most audio-visual deepfake detectors rely on lip–speech synchronization and are designed for static, frontal videos, limiting their reliability in dynamic, in-the-wild recordings where the speaker moves and lip cues are degraded or missing. As an alternative, we propose a physics-aware detector for dynamic speaking videos that leverages an acoustic constraint: in real recordings, speech energy varies predictably with speaker-to-camera distance, whereas deepfake generation can break this coupling. Our method combines distance estimates from video with distance-dependent acoustic measures from speech to detect deepfakes. Experiments on in-the-wild benchmarks show that the proposed physical cue is effective in dynamic scenarios, and that it complements lip-sync-based detectors, with a simple ensemble achieving consistent gains across datasets.

Speech-Distance Physical Consistency

Methodology Framework

In authentic recordings, the speech signal obeys physical laws of acoustic attenuation. Specifically, the direct sound pressure inversely correlates with the distance between the speaker and the microphone. Our framework models this physical consistency by extracting visual depth/distance cues and aligning them with acoustic metrics such as Signal-to-Noise Ratio (SNR) and Clarity ($C_{50}$) to capture acoustic features. Deepfake models often fail to satisfy these multi-modal physical boundaries, allowing for robust out-of-domain detection.

BibTeX

@InProceedings{kim2026physics,
  author    = {Kim, Kyeongrae and Kim, Sung-Bin and Oh, Hyun-Bin and Oh, Tae-Hyun},
  title     = {Physics-Aware Deepfake Detection via Speech-Distance Consistency},
  booktitle = {Proceedings of Interspeech},
  year      = {2026}
}