You are here

ImageCLEF MEDVQA-GI

Motivation

The MEDVQA-GI challenge, now in its fourth iteration, continues to advance research in Visual Question Answering (VQA) for gastrointestinal (GI) endoscopy with a strong emphasis on clinical relevance, explainability, and safety. Building on insights and outcomes from previous editions, this year's challenge shifts focus from data generation toward trustworthy multimodal reasoning in realistic clinical settings.

The 2026 challenge emphasizes models that not only answer clinically relevant questions from GI endoscopy images, but also justify their answers in a manner aligned with medical reasoning while adhering to clinical safety principles. In addition, the challenge explicitly addresses emerging concerns related to overconfidence, hallucinations, and misleading explanations in medical AI systems. By integrating behavioral safety evaluation and retrieval-augmented reasoning, the challenge aims to promote the development of reliable, interpretable, and clinically robust VQA systems for endoscopy.

Task Description

We define two subtasks for this year's challenge.

Subtask 1: Clinically Relevant Visual Question Answering

This subtask asks participants to develop algorithms that accurately answer clinically relevant questions based on gastrointestinal (GI) endoscopy images. Models must correctly interpret visual findings and question intent to produce concise and medically sound answers. The focus is on diagnostic and descriptive questions commonly encountered in clinical practice.

Subtask 2: Multimodal Explainability and Safety-Aware Reasoning

This subtask extends standard VQA by requiring participants to generate coherent multimodal justifications that combine textual explanations with visual evidence from the image. Explanations should reflect established medical reasoning and remain factual, consistent, and clinically safe. A dedicated safety layer evaluates whether models exhibit undesirable behaviors such as unwarranted overconfidence, misleading explanations, or deviations from medical best practices.

Data

The data used for this year's challenge is based on an expanded version of the Kvasir-VQA-x1 dataset, comprising more than 150,000 question-answer pairs derived from GI endoscopy images. The dataset supports both answer prediction and explainability evaluation.

To facilitate retrieval-augmented generation approaches, participants are also provided with a curated collection of verified endoscopy-related clinical resources that may be used during inference.

Datasets will be made available on our GitHub:
https://github.com/simula/ImageCLEFmed-MEDVQA-GI-2026

Evaluation methodology

The evaluation will differ for each task and will include a combination of objective and expert-based assessments. Subtask 1 is evaluated using automatic metrics assessing answer correctness and language quality. Subtask 2 combines automatic measures with expert-based assessment of explanation quality, interpretability, clinical plausibility, and safety.

A dedicated behavioral safety evaluation examines whether generated answers and explanations exhibit problematic behaviors such as hallucination, overconfidence, or unsafe clinical recommendations. We will host an online leaderboard for rapid evaluation feedback to nurture a competitive environment.

Participant registration

Please refer to the general
ImageCLEF registration instructions.

Please also email steven@simula.no to register your interest.

Note: Thie task does not use the Ai4Media benchmarking platform for evaluation. Creating an account on the platform for this task is not required. Completing the registration form is still required, similar to the other tasks. You can also find the form here.

Preliminary Schedule

  • 26.01.2026: Registration opens for all ImageCLEF tasks
  • 01.03.2026: Development dataset released
  • 15.03.2026: Test dataset released
  • 23.04.2026: Registration closes for all ImageCLEF tasks
  • 07.05.2026: Deadline for submitting participant runs
  • 14.05.2026: Release of the processed results by the task organizers (depends on task)
  • 28.05.2026: Submission of participant papers [CEUR-WS]
  • 30.06.2026: Notification of acceptance
  • 21.09.2026: CLEF 2026, Jena, Germany

Submission Instructions

Please refer to our
GitHub page
for detailed submission instructions.

Results

TBA

CEUR Working Notes

For detailed instructions, please refer to
this PDF file.
A summary of the most important points:

  • All participating teams with at least one graded submission, regardless of the score, should submit a CEUR working notes paper.
  • Teams who participated in both subtasks should generally submit only one report.
  • You can find the working notes template here: https://clef2026.clef-initiative.eu/calls/submitting

Citations

When referring to MedVQA, pelase cite the following:

@inproceedings{ImageCLEFmedicalVQAOverview2026,
title = {Overview of ImageCLEFmedical 2026 – Medical Visual Question Answering for Gastrointestinal Tract},
author = {Hicks, Steven A. and Gautam, Sushant and Halvorsen, Pål and Riegler, Michael A. and Thambawita, Vajira},
year = 2026,
month = {September},
booktitle = {CLEF2026 Working Notes},
publisher = {CEUR-WS.org},
address = {Berlin, Germany},
series = {{CEUR} Workshop Proceedings}
}

When referring to ImageCLEF 2026, please cite the following:

@inproceedings{ImageCLEF2026,
title = {Overview of ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security},
author = {Bogdan Ionescu and Henning M{\"u}ller and Dan{-}Cristian Stanciu and Andrei Radu and Radu{-}George Bolborici and Marian Negru and Alexandru{-}Florin Ene and Vlad{-}Mihai Vasilescu and Ana-Antonia Nicolae and Liviu{-}Daniel \c{S}tefan and Mihai{-}Gabriel Constantin and Mihai Dogariu and Alexandra{-}Georgiana Andrei and Hendrik Damm and Tabea M. G. Pakull and Asma {Ben Abacha} and Alba {Garc\'ia Seco de Herrera} and Christoph M. Friedrich and Raphael Br{\"u}ngel and Lea Reinartz and Henning Sch{\"a}fer and Cynthia Sabrina Schmidt and Benjamin Bracke and Praveen Nath and Bahad{\i}r Ery{\i}lmaz and Maja Hjuler and Diandra Fabre and Claire Lemaire and Benjamin Lecouteux and Didier Schwab and Dimitar Dimitrov and Ming Shan Hee and Momina Ahsan and Sarfraz Ahmad and Dimitrina Zlatkova and Georgi Pachov and Zhuohan Xie and Preslav Nakov and Ivan Koychev and Juampablo E. {Heras Rivera} and Daniel K. Low and Wen{-}wai Yim and Jacob Ruzevick and Dan Child and Mehmet Kurt and Zhaoyi Sun and Fei Xia and Meliha Yetisgen and Ahmedkhan Radzhabov and Yuri Prokopchuk and Vassili Kovalev and Dzmitry Karpenka and Steven A. Hicks and Sushant Gautam and Michael A. Riegler and Vajira Thambawita and P\r{a}l Halvorsen and Mohammad {El Sakka} and Josiane Mothe and Alexandra B\u{a}icoianu and Corneliu{-}Nicolae Florea and Mihai Ivanovici},
booktitle = {Experimental IR Meets Multilinguality, Multimodality, and Interaction},
series = {Proceedings of the Seventeenth International Conference of the CLEF Association (CLEF 2026)},
year = {2026},
month = {September 21--24},
address = {Jena, Germany},
publisher = {Springer Lecture Notes in Computer Science LNCS},
}

Contact

Organizers: