All papers
Deepfakes & securityFull-paper digest

MS-GNN: Multi-Scale Graph Neural Network for Detecting Local Audio-Visual Forgery Traces

Jianrong Wang, Hengyang Guo, Jie Liu, Ju Zhang, Qi Li, Ying Guo, Jing Zhao

Explore this paper with your agent

Get a guided paper analysis: method, diagrams, experiments, and metrics. Includes the full wiki digest and metadata.

13.2 KB · Ready to paste

Preview copied content

TL;DR — MS-GNN is a multi-scale graph neural network designed to detect localized audio-visual deepfakes by preventing transient manipulation cues from being diluted by global pooling. It achieves 98.12% accuracy on the LAV-DF dataset.

Key contributions

  • Formulates audio-visual forgery detection as a multi-scale graph problem, using non-overlapping temporal windows to isolate short-span tampering traces.
  • Proposes a bi-directional multi-scale graph architecture featuring bottom-up semantic aggregation and top-down context-guided refinement.
  • Introduces time-aware attention with a Gaussian temporal bias and Dual Cross-modal Graph Attention with learnable identity gates to suppress untrained cross-modal noise.
  • Employs a multi-task training objective combining video-level and auxiliary window-level focal losses to preserve both fine-grained localization and global discrimination.

Problem

Modern deepfake generation tools often restrict manipulations to brief, specific temporal windows (e.g., replacing a single spoken phrase or injecting a localized lip-sync error) while leaving the remainder of the video authentic. Traditional audio-visual detectors rely on global feature pooling, which lowers the signal-to-noise ratio by treating local tampering evidence as background noise and leading to missed detections. Prior paradigms like DiMoDif and Referee rely heavily on high-level semantic discourse or external reference clips, leaving them vulnerable to non-semantic artifacts in reference-free settings.

Method

The framework partitions input video streams into non-overlapping 400 ms windows to match natural word durations. Visual features (768-dim) from TimeSformer and audio features (768-dim) from WavLM are concatenated with temporal encodings and processed through a heterogeneous graph where nodes represent audio and visual windows, and edges model temporal adjacency (with a time-aware Gaussian attention bias parameter τ\tau) and cross-modal synchronization.

Each MS-GNN layer operates in three stages: (1) Window-level interaction via dual-stream graph attention with a learnable scalar identity gate γ\gamma initialized near zero (e.g., -0.045, 0.058, 0.012) to protect pre-trained uni-modal features; (2) Bottom-up semantic aggregation, grouping windows into overlapping segments of size Ssize=8S_{size}=8 (3.2 seconds) processed via Transformers with CLS tokens and cross-attention; and (3) Top-down contextual refinement, injecting averaged segment representations back into the window nodes via residual connections.

The final window representations pass through attention pooling to produce a global video vector. The model is trained for 30 epochs using the AdamW optimizer (learning rate 1×10−41\times 10^{-4}, batch size 256, CosineAnnealingLR) with a multi-task focal loss objective combining video-level (λvid=0.9\lambda_{vid}=0.9) and window-level (λwin=0.1\lambda_{win}=0.1) supervision.

Experimental setup

Evaluated on LAV-DF (rich in local forgeries) and FakeAVCeleb (large-scale, mostly global manipulations evaluated via 5-fold cross-validation). Compared against baselines including MDS, JAVDD, BA-TFD, DimoDif, Referee, VFD, AVOID-DF, and MRDF-CE using Accuracy (ACC) and Area Under ROC Curve (AUC). Notable implementation details include an NVIDIA A800 GPU for complexity tracking, 3 MS-GNN layers with hidden dimensions [512, 256, 64], and TimeSformer/WavLM frozen/extracted features.

Results

On LAV-DF, MS-GNN achieves 98.12% ACC and 99.81% AUC, outperforming DimoDif (97.84% ACC) and Referee (96.74% ACC). On FakeAVCeleb, which lacks isolated local anomalies and favors global models, it records a competitive 93.20% ACC and the highest AUC of 96.75%. Ablation studies confirm that combining window interaction and hierarchical segment fusion yields a synergistic +6.32% ACC gain over a base graph network (91.80% to 98.12%), and demonstrate that the multi-scale graph architecture remains highly effective (98.01% ACC) even when auxiliary window-level supervision is omitted.

SystemLAV-DF ACC (%)LAV-DF AUC (%)FakeAVCeleb ACC (%)FakeAVCeleb AUC (%)
MDS / VFD93.4297.8381.5286.11
DimoDif97.8499.74--
Referee96.7499.7083.7589.86
MRDF--94.0592.43
MS-GNN (ours)98.1299.8193.2096.75

Limitations

The model's end-to-end latency is bottlenecked by heavy offline feature extractors (TimeSformer and WavLM), requiring over 2.3 seconds per 2.8-second clip despite the lightweight MS-GNN graph module running at 0.123 G FLOPs. Performance relies on fixed window sizes (400 ms) and segment counts, which may require tuning for non-standard speech rates or music-heavy videos. The evaluation is currently restricted to standard English-centric or benchmark multi-modal datasets, leaving multilingual and in-the-wild generalization unproven.

Why read this

Read this paper if you work on multi-modal deepfake detection and need a principled hierarchical graph approach to resolve the SNR drop caused by global pooling. It provides a blueprint for combining window-level local temporal tracking with global context without relying on external reference videos.

Code

None released (as of this page's updated date). If you are an author with a repo, please claim this entry — see CONTRIBUTING.md.

Applications

Automated multi-modal deepfake video screening, forensic verification of news media, and localized audio-visual tampering localization.

Institutions

Tianjin University, Tianjin Renai College, Tianjin Electronic Information College, Tianjin University of Technology, Tianjin Beiyang Rongke Intelligent Technology Co., Ltd

Funding / 經費: Key R&D Program of the Nanning Science Research and Technology Development Plan, Tianjin Science and Technology Program, Special Project for High-Quality Development of Manufacturing Industry

All 950k paper pairs scored by TypeSafe Jev (scripts/related/); relatedness 0 = unrelated … 3 = directly comparable.

SOURCE & COVERAGE

AI-assisted full-paper digest. Check important claims against the original paper.

DOI: 10.21437/Interspeech.2026-416