CERESResearch Repository

VIDS-guard: a novel forensics-aware multi-stream transformer framework for robust deepfake video detection

Loading...
Thumbnail Image

Date published

Free to read from

2026-05-28

Supervisor/s

Industry supervisor/s

Journal Title

Journal ISSN

Volume Title

Publisher

Department

Course name

ISSN

2667-3053

Format

Citation

Alanazi S, Asif S. (2026) VIDS-guard: a novel forensics-aware multi-stream transformer framework for robust deepfake video detection. Intelligent Systems with Applications, Volume 30, May 2026, Article number 200664

Abstract

The proliferation of highly realistic deepfake videos poses a growing threat to digital trust, underscoring the need for detectors that remain reliable across diverse manipulation types and capture conditions. This paper introduces VIDS-Guard (Video Integrity Deepfake Shield), a novel forensics-aware multi-stream transformer framework that integrates spatial, frequency, and temporal cues within a unified architecture. Unlike conventional convolutional or transformer-based detectors that rely primarily on semantic consistency, VIDS-Guard embeds forensic inductive biases through Spatial Rich Model (SRM) residual filtering, YCbCr color-space decomposition, and Fast Fourier Transform (FFT) spectral embeddings to expose subtle manipulation artifacts. A temporal transformer encoder with attention pooling further models cross-frame inconsistencies, enabling robust video-level predictions. Extensive experiments conducted with six benchmark models—Xception, ResNet50, MobileNetV3-Large, SlowFast, ViViT, and TimeSformer—demonstrate that VIDS-Guard achieves superior generalization and balanced detection performance across validation, test, and unseen datasets, attaining the highest accuracy and Macro-F1 scores under domain shift. These findings establish VIDS-Guard as a state-of-the-art forensic framework for trustworthy multimedia authentication and emphasize the importance of incorporating forensic priors to ensure sustainable robustness in deepfake video detection.

Description

Software description

Software language

Git repository

Keywords

46 Information and Computing Sciences, 40 Engineering, 4008 Electrical Engineering, 4603 Computer Vision and Multimedia Computation, Deepfake detection, Video forensics, Temporal modeling, Frequency embedding, Transformer architecture, Multimedia security

DOI

Rights

Attribution 4.0 International

Funder/s

This work was supported by the Engineering and Physical Sciences Research Council (EPSRC) as part of “Made Smarter Innovation - Research Centre for Smart, Collaborative Industrial Robotics” [grant number EP/V062158/1]

Grant number

Relationships

Resources