Speech-Distance Physical Consistency
In authentic recordings, the speech signal obeys physical laws of acoustic attenuation. Specifically, the direct sound pressure inversely correlates with the distance between the speaker and the microphone. Our framework models this physical consistency by extracting visual depth/distance cues and aligning them with acoustic metrics such as Signal-to-Noise Ratio (SNR) and Clarity ($C_{50}$) to capture acoustic features. Deepfake models often fail to satisfy these multi-modal physical boundaries, allowing for robust out-of-domain detection.