CERESResearch Repository

Multimodal misinformation detection across diverse languages using RAG and LLMs

dc.contributor.authorHarris, Sheetal
dc.contributor.authorTa, Vinh Thong
dc.contributor.authorTrovati, Marcello
dc.contributor.authorNakhla, Ghada
dc.contributor.authorLatif, Faiza
dc.contributor.authorKorkontzelos, Ioannis
dc.date.accessioned2026-04-30T08:49:58Z
dc.date.available2026-04-30T08:49:58Z
dc.date.freetoread2026-04-30
dc.date.issued2026-12-31
dc.date.pubOnline2026-04-09
dc.description.abstractThe rapid spread of multimodal fake news (FN) on Online Social Networks (OSNs) threatens digital information ecosystems, particularly in low-resource languages. Existing multimodal fake news detection (FND) methods are largely limited to high-resource settings, restricting their global applicability. We propose an M&M-RAG, a Multilingual & Multimodal Retrieval-Augmented Generation framework, that leverages Large Vision-Language Models (LVLMs) and Large Language Models (LLMs) to verify news claims across English, Chinese and Urdu. M&M-RAG integrates real-time multilingual evidence retrieval, language-aware prompting, and cross-modal reasoning for fact verification. We further propose Multi-Ax-to-Grind Urdu, the first large-scale, multi-domain multimodal benchmark for FND in Urdu. Experiments on typologically diverse monolingual multimodal datasets demonstrate that M&M-RAG achieves state-of-the-art (SOTA) performance, with 94.6% accuracy and 94.2% F1 score, surpassing models such as SpotFake, MPFN, MMCFND, and Semi-FND. The proposed framework remains robust in zero-shot and cross-lingual scenarios under frozen-model inference without task-specific fine-tuning. The results underscore the scalability and interpretability of LVLM-based approaches for combating multimodal misinformation, particularly in under-represented and typologically diverse languages.
dc.description.journalNameJournal of Intelligent Information Systems
dc.description.sponsorshipThe last author has participated in this research work as part of ALFIE Project, which has received funding by the European Union’s Horizon Europe research and innovation programme, under Grant Agreement No. 101177912.
dc.identifier.citationHarris S, Ta VT, Trovati M, Nakhla G, et al., (2026) Multimodal misinformation detection across diverse languages using RAG and LLMs. Journal of Intelligent Information Systems, Available online 9 April 2026en_UK
dc.identifier.eissn1573-7675
dc.identifier.elementsID870277
dc.identifier.issn0925-9902
dc.identifier.urihttps://doi.org/10.1007/s10844-026-01042-x
dc.identifier.urihttps://dspace.lib.cranfield.ac.uk/handle/1826/25193
dc.languageEnglish
dc.language.isoen
dc.publisherSpringeren_UK
dc.publisher.urihttps://link.springer.com/article/10.1007/s10844-026-01042-x
dc.relation.isreferencedbyhttps://figshare.com/s/62b9bbda2464d2059eeb
dc.rightsAttribution 4.0 Internationalen
dc.rights.urihttp://creativecommons.org/licenses/by/4.0/
dc.subjectMultimodal multilingual fake news detectionen_UK
dc.subjectLarge vision-language models (LVLMs)en_UK
dc.subjectRetrieval-augmented generation (RAG)en_UK
dc.subjectNLPen_UK
dc.subject4605 Data Management and Data Scienceen_UK
dc.subjectSocial Determinants of Healthen_UK
dc.subjectInformation Systemsen_UK
dc.subject46 Information and computing sciencesen_UK
dc.titleMultimodal misinformation detection across diverse languages using RAG and LLMsen_UK
dc.typeArticle
dc.type.subtypeJournal Article
dcterms.dateAccepted2026-03-20

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Multimodal_misinformation_detection-2026.pdf
Size:
2.9 MB
Format:
Adobe Portable Document Format
Description:
Published version

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
1.63 KB
Format:
Plain Text
Description: