VIDS-guard: a novel forensics-aware multi-stream transformer framework for robust deepfake video detection
| dc.contributor.author | Alanazi, Sami | |
| dc.contributor.author | Asif, Seemal | |
| dc.date.accessioned | 2026-05-28T10:24:20Z | |
| dc.date.available | 2026-05-28T10:24:20Z | |
| dc.date.freetoread | 2026-05-28 | |
| dc.date.issued | 2026-05 | |
| dc.date.pubOnline | 2026-04-13 | |
| dc.description.abstract | The proliferation of highly realistic deepfake videos poses a growing threat to digital trust, underscoring the need for detectors that remain reliable across diverse manipulation types and capture conditions. This paper introduces VIDS-Guard (Video Integrity Deepfake Shield), a novel forensics-aware multi-stream transformer framework that integrates spatial, frequency, and temporal cues within a unified architecture. Unlike conventional convolutional or transformer-based detectors that rely primarily on semantic consistency, VIDS-Guard embeds forensic inductive biases through Spatial Rich Model (SRM) residual filtering, YCbCr color-space decomposition, and Fast Fourier Transform (FFT) spectral embeddings to expose subtle manipulation artifacts. A temporal transformer encoder with attention pooling further models cross-frame inconsistencies, enabling robust video-level predictions. Extensive experiments conducted with six benchmark models—Xception, ResNet50, MobileNetV3-Large, SlowFast, ViViT, and TimeSformer—demonstrate that VIDS-Guard achieves superior generalization and balanced detection performance across validation, test, and unseen datasets, attaining the highest accuracy and Macro-F1 scores under domain shift. These findings establish VIDS-Guard as a state-of-the-art forensic framework for trustworthy multimedia authentication and emphasize the importance of incorporating forensic priors to ensure sustainable robustness in deepfake video detection. | |
| dc.description.journalName | Intelligent Systems with Applications | |
| dc.description.sponsorship | This work was supported by the Engineering and Physical Sciences Research Council (EPSRC) as part of “Made Smarter Innovation - Research Centre for Smart, Collaborative Industrial Robotics” [grant number EP/V062158/1] | |
| dc.identifier.citation | Alanazi S, Asif S. (2026) VIDS-guard: a novel forensics-aware multi-stream transformer framework for robust deepfake video detection. Intelligent Systems with Applications, Volume 30, May 2026, Article number 200664 | en_UK |
| dc.identifier.elementsID | 870209 | |
| dc.identifier.issn | 2667-3053 | |
| dc.identifier.paperNo | 200664 | |
| dc.identifier.uri | https://doi.org/10.1016/j.iswa.2026.200664 | |
| dc.identifier.uri | https://dspace.lib.cranfield.ac.uk/handle/1826/25213 | |
| dc.identifier.volumeNo | 30 | |
| dc.language | English | |
| dc.language.iso | en | |
| dc.publisher | Elsevier | en_UK |
| dc.publisher.uri | https://www.sciencedirect.com/science/article/pii/S2667305326000396?via%3Dihub | |
| dc.relation.isreferencedby | https://github.com/IFRA-Cranfield/VIDS-Guard | |
| dc.relation.isreferencedby | https://doi.org/10.5281/zenodo.17362749 | |
| dc.relation.isreferencedby | https://doi.org/10.5281/zenodo.17382113 | |
| dc.rights | Attribution 4.0 International | en |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | 46 Information and Computing Sciences | en_UK |
| dc.subject | 40 Engineering | en_UK |
| dc.subject | 4008 Electrical Engineering | en_UK |
| dc.subject | 4603 Computer Vision and Multimedia Computation | en_UK |
| dc.subject | Deepfake detection | en_UK |
| dc.subject | Video forensics | en_UK |
| dc.subject | Temporal modeling | en_UK |
| dc.subject | Frequency embedding | en_UK |
| dc.subject | Transformer architecture | en_UK |
| dc.subject | Multimedia security | en_UK |
| dc.title | VIDS-guard: a novel forensics-aware multi-stream transformer framework for robust deepfake video detection | en_UK |
| dc.type | Article | |
| dc.type.subtype | Journal Article | |
| dcterms.dateAccepted | 2026-04-08 |
