You are here

MultimodalReasoning

Motivation

Vision-Language Models (VLMs) have demonstrated remarkable progress in integrating visual and textual information, achieving strong performance in tasks such as image captioning, visual question answering (VQA), and multimodal dialogue. Despite these advancements, their ability to perform structured reasoning and draw inferences from complex visual-linguistic relationships remains limited. In particular, VLMs often struggle with questions that demand multi-step reasoning, abstract understanding, or hypothetical thinking grounded in visual evidence. This year, we plan to expand the existing set of multiple-choice questions and introduce a new task designed to challenge VLMs’ reasoning abilities even further. The study of visual question answering and visual reasoning thus provides a crucial benchmark for evaluating how effectively modern models can understand, interpret, and reason over multimodal inputs presented across diverse domains and languages.

News

This year we are using a new website for task information: https://mbzuai-nlp.github.io/ImageCLEF-MultimodalReasoning/2026/

The Test data from ImageCLEF 2025 with released labels - https://huggingface.co/datasets/MBZUAI/EXAMS-V. This repository also contains the training and development data.

GitHub: https://github.com/mbzuai-nlp/ImageCLEF-2025-MultimodalReasoning, where you will find information about the previous edition of the task. We will update the repository accordingly with new instructions and scripts.

CEUR Working Notes

Paper Submission Instructions

The full schedule is available at:
https://clef2026.clef-initiative.eu/dates/

Important Dates

  • End of evaluation cycle (submission of runs): 7 May 2026
  • Submission of participant papers (CEUR-WS): 28 May 2026
  • Notification of acceptance for participant papers (CEUR-WS): 30 June 2026
  • Camera-ready submission of participant papers: 6 July 2026

All submissions, reviews, and camera-ready versions will be handled through EasyChair: EasyChair CLEF 2026
A separate EasyChair track will be created for each lab/workshop. Please make sure you submit in the correct track.
The papers will go through a review process and will receive a decision from the lab organizers.

The participant papers should be written using the template provided here:
CLEF 2026 Working Notes Submission Template

Submissions are expected to be in English language and 5 pages minimum, with no maximum page limit.

Citations

When referring to ImageCLEF 2026 Overview of the Multimodal Reasoning task, please cite the following:

@inproceedings{ImageCLEFMultimodalReasoningTaskOverview2026,
title = {{O}verview of the {I}mage{CLEF} 2026 {T}ask on {M}ultimodal {R}easoning},
author = {Dimitrov, Dimitar and Hee, Ming Shan and Ahsan, Momina and Ahmad, Sarfraz and Zlatkova, Dimitrina and Pachov, Georgi and Xie, Zhuohan and Nakov, Preslav and Koychev, Ivan},
booktitle = {CLEF 2026 Working Notes},
series = {CEUR Workshop Proceedings},
year = {2026},
month = {September 21--24},
address = {Jena, Germany},
publisher = {CEUR-WS.org},
}
When referring to ImageCLEF 2026, please cite the following:

@inproceedings{ImageCLEF2026,
title = {Overview of ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security},
author = {Bogdan Ionescu and Henning M{\"u}ller and Dan{-}Cristian Stanciu and Andrei Radu and Radu{-}George Bolborici and Marian Negru and Alexandru{-}Florin Ene and Vlad{-}Mihai Vasilescu and Ana-Antonia Nicolae and Liviu{-}Daniel \c{S}tefan and Mihai{-}Gabriel Constantin and Mihai Dogariu and Alexandra{-}Georgiana Andrei and Hendrik Damm and Tabea M. G. Pakull and Asma {Ben Abacha} and Alba {Garc\'ia Seco de Herrera} and Christoph M. Friedrich and Raphael Br{\"u}ngel and Lea Reinartz and Henning Sch{\"a}fer and Cynthia Sabrina Schmidt and Benjamin Bracke and Praveen Nath and Bahad{\i}r Ery{\i}lmaz and Maja Hjuler and Diandra Fabre and Claire Lemaire and Benjamin Lecouteux and Didier Schwab and Dimitar Dimitrov and Ming Shan Hee and Momina Ahsan and Sarfraz Ahmad and Dimitrina Zlatkova and Georgi Pachov and Zhuohan Xie and Preslav Nakov and Ivan Koychev and Juampablo E. {Heras Rivera} and Daniel K. Low and Wen{-}wai Yim and Jacob Ruzevick and Dan Child and Mehmet Kurt and Zhaoyi Sun and Fei Xia and Meliha Yetisgen and Ahmedkhan Radzhabov and Yuri Prokopchuk and Vassili Kovalev and Dzmitry Karpenka and Steven A. Hicks and Sushant Gautam and Michael A. Riegler and Vajira Thambawita and P\r{a}l Halvorsen and Mohammad {El Sakka} and Josiane Mothe and Alexandra B\u{a}icoianu and Corneliu{-}Nicolae Florea and Mihai Ivanovici},
booktitle = {Experimental IR Meets Multilinguality, Multimodality, and Interaction},
series = {Proceedings of the Seventeenth International Conference of the CLEF Association (CLEF 2026)},
year = {2026},
month = {September 21--24},
address = {Jena, Germany},
publisher = {Springer Lecture Notes in Computer Science LNCS},
}