Multimodal misinformation detection across diverse languages using RAG and LLMs
| dc.contributor.author | Harris, Sheetal | |
| dc.contributor.author | Ta, Vinh Thong | |
| dc.contributor.author | Trovati, Marcello | |
| dc.contributor.author | Nakhla, Ghada | |
| dc.contributor.author | Latif, Faiza | |
| dc.contributor.author | Korkontzelos, Ioannis | |
| dc.date.accessioned | 2026-04-30T08:49:58Z | |
| dc.date.available | 2026-04-30T08:49:58Z | |
| dc.date.freetoread | 2026-04-30 | |
| dc.date.issued | 2026-12-31 | |
| dc.date.pubOnline | 2026-04-09 | |
| dc.description.abstract | The rapid spread of multimodal fake news (FN) on Online Social Networks (OSNs) threatens digital information ecosystems, particularly in low-resource languages. Existing multimodal fake news detection (FND) methods are largely limited to high-resource settings, restricting their global applicability. We propose an M&M-RAG, a Multilingual & Multimodal Retrieval-Augmented Generation framework, that leverages Large Vision-Language Models (LVLMs) and Large Language Models (LLMs) to verify news claims across English, Chinese and Urdu. M&M-RAG integrates real-time multilingual evidence retrieval, language-aware prompting, and cross-modal reasoning for fact verification. We further propose Multi-Ax-to-Grind Urdu, the first large-scale, multi-domain multimodal benchmark for FND in Urdu. Experiments on typologically diverse monolingual multimodal datasets demonstrate that M&M-RAG achieves state-of-the-art (SOTA) performance, with 94.6% accuracy and 94.2% F1 score, surpassing models such as SpotFake, MPFN, MMCFND, and Semi-FND. The proposed framework remains robust in zero-shot and cross-lingual scenarios under frozen-model inference without task-specific fine-tuning. The results underscore the scalability and interpretability of LVLM-based approaches for combating multimodal misinformation, particularly in under-represented and typologically diverse languages. | |
| dc.description.journalName | Journal of Intelligent Information Systems | |
| dc.description.sponsorship | The last author has participated in this research work as part of ALFIE Project, which has received funding by the European Union’s Horizon Europe research and innovation programme, under Grant Agreement No. 101177912. | |
| dc.identifier.citation | Harris S, Ta VT, Trovati M, Nakhla G, et al., (2026) Multimodal misinformation detection across diverse languages using RAG and LLMs. Journal of Intelligent Information Systems, Available online 9 April 2026 | en_UK |
| dc.identifier.eissn | 1573-7675 | |
| dc.identifier.elementsID | 870277 | |
| dc.identifier.issn | 0925-9902 | |
| dc.identifier.uri | https://doi.org/10.1007/s10844-026-01042-x | |
| dc.identifier.uri | https://dspace.lib.cranfield.ac.uk/handle/1826/25193 | |
| dc.language | English | |
| dc.language.iso | en | |
| dc.publisher | Springer | en_UK |
| dc.publisher.uri | https://link.springer.com/article/10.1007/s10844-026-01042-x | |
| dc.relation.isreferencedby | https://figshare.com/s/62b9bbda2464d2059eeb | |
| dc.rights | Attribution 4.0 International | en |
| dc.rights.uri | http://creativecommons.org/licenses/by/4.0/ | |
| dc.subject | Multimodal multilingual fake news detection | en_UK |
| dc.subject | Large vision-language models (LVLMs) | en_UK |
| dc.subject | Retrieval-augmented generation (RAG) | en_UK |
| dc.subject | NLP | en_UK |
| dc.subject | 4605 Data Management and Data Science | en_UK |
| dc.subject | Social Determinants of Health | en_UK |
| dc.subject | Information Systems | en_UK |
| dc.subject | 46 Information and computing sciences | en_UK |
| dc.title | Multimodal misinformation detection across diverse languages using RAG and LLMs | en_UK |
| dc.type | Article | |
| dc.type.subtype | Journal Article | |
| dcterms.dateAccepted | 2026-03-20 |
