Welcome to the 4th edition of the GANs task!
Motivation

AI and computer-aided diagnosis systems for medical tasks such as prediction, detection, and classification depend on access to large and diverse datasets for effective training. High-quality data enable these models to learn complex patterns, improving both accuracy and reliability. However, obtaining real medical data is difficult due to strict privacy regulations and ethical concerns. Patients typically consent to the use of their medical information for their own clinical care, but not for broad research purposes. As a result, collecting sufficiently large datasets for training AI models remains a significant challenge and slows progress in developing advanced healthcare tools.
Synthetic medical data has emerged as a promising solution. Generative models such as GANs (Generative Adversarial Networks) can produce realistic-looking data that resemble true medical images while not being tied to any specific patient. These synthetic datasets allow researchers to build and evaluate AI systems without directly relying on sensitive patient information, potentially easing data-access barriers and enabling greater diversity and scalability.
However, an important concern remains: synthetic data must be free of hidden traces or “fingerprints” from the real images used during training. If synthetic images inadvertently reproduce patient-specific details, they could still leak private information. Ensuring that generative models do not memorize or reveal identifiable patterns is therefore crucial. Only truly privacy-preserving synthetic data can safely accelerate innovation in AI-powered healthcare.
Lessons learned
-
In the first three editions of this task (2023-2025), various generative models were analyzed within the framework of the first subtask to investigate whether synthetic images contained "fingerprints" of the real medical data used during training. The results demonstrated that generative models do retain and imprint features from their training data, raising important security and privacy concerns. These findings underscore the need for robust techniques to detect and mitigate such imprints to ensure that synthetic images protect patient privacy while maintaining their utility for research and development.
-
In the 2nd edition of the task (2024), it was confirmed that generative models leave unique "fingerprints" on the synthetic images they produce. By analyzing images generated from various models, distinct patterns and features were identified that allowed the attribution of synthetic images to their respective generative models.
-
In the 3rd edition (2025), subtask “Identify Training Data Subsets” explored attribution at a broader scale: identifying the training subset used to generate synthetic images from a diffusion model. Identification performance was significantly higher, with multiple teams exceeding 98% accuracy. While this task did not involve image-level membership inference, it revealed that diffusion-generated images can still reflect statistical characteristics of the training data subsets. The success of various supervised and semi-supervised classification techniques suggests that diffusion models, while effective at preserving image realism, may still encode latent information about their source data distributions. This raises important considerations about the use of synthetic data for data sharing or augmentation, particularly when the provenance or diversity of training data must remain confidential.
News
Preliminary Schedule
- 26.01.2026: Registration opens for all ImageCLEF tasks
- 02.02.2026: Development dataset released (depends on task)
- 09.03.2026: Test dataset released (depends on task)
- 23.04.2026: Registration closes for all ImageCLEF tasks
0710.05.2026: Deadline for submitting participant runs
1415.05.2026: Release of the processed results by the task organizers (depends on task)
- 28.05.2026: Submission of participant papers [CEUR-WS]
- 30.06.2026: Notification of acceptance
- 21.09.2026: CLEF 2026, Jena, Germany
Task Description
In 2025, we continue the traditional ImageCLEFmed GAN challenge task on detecting whether real medical images were used to train generative models. Building on this foundation, we introduce two new subtasks that explore latent-space privacy leakage and the generation of privacy-preserving CT images. Together, these subtasks aim to deepen our understanding of how generative models memorize, reconstruct, or unintentionally expose sensitive biomedical information.
Subtask 1: Detect training data usage
In this subtask, participants will analyze synthetic biomedical images to determine whether specific real images were used in the training process of generative models. For each real image in the test set, participants must assign a binary label: 1 if the image was used during training or 0 if the image was not used.
The goal is to detect the presence of training-data "fingerprints" within synthetic outputs and to assess vulnerabilities related to memorization.
Subtask 2: Identify corresponding latent vectors
In this subtask, participants will analyze a set of synthetic images and a collection of latent vectors (128-dimensional, sampled from a standard normal distribution) and determine which latent vector corresponds to which generated image. This task evaluates how faithfully generative models map latent vectors to output images and whether this mapping reveals structural or patient-specific information.
A well-generalizing model should produce synthetic images that do not reflect identifiable traces of specific training samples, even when latent vectors are disclosed. In contrast, an overfitted model may encode training-data fingerprints within the latent space, enabling reconstruction or linkage back to individuals.
By requiring participants to match images to latent vectors, this subtask measures the predictability of the latent-to-image mapping, the presence of hidden structure or memorization within the latent space, and the robustness of generative models to latent-space–based privacy attacks. This subtask therefore complements Subtask 1 by examining privacy leakage not only through image similarity, but through learned representations.
Subtask 3: Privacy-preserving CT slice generation
Generative models trained on medical data may unintentionally memorize and reproduce sensitive patient information. In this subtask, participants are asked to generate new CT slices using a set of real slices as guidance. Although the task resembles a typical image-generation problem, its primary goal is to study privacy leakage and fingerprint propagation by studying the following hypotheses: models that overfit may reconstruct patient-specific anatomy or replicate identifiable features; (ii) models that generalize well should produce realistic CT images without encoding individual patient identities.
This subtask allows us to quantify the trade-off between realism and privacy preservation. Submitted images will be evaluated for fidelity, diversity, and the degree to which sensitive training-data fingerprints are preserved or eliminated.
Data
The benchmarking dataset includes both real and synthetic biomedical images. The real images consist of axial slices of 3D CT scans from approximately 8,000 lung tuberculosis patients. These slices vary in appearance: some may look relatively “normal,” while others exhibit distinct lung lesions, including severe cases. The real images are stored in 8-bit per pixel PNG format, with dimensions of 256x256 pixels, providing a standardized resolution for analysis.
The synthetic images, also sized at 256x256 pixels, have been generated using various generative models, including Generative Adversarial Networks (GANs) and Diffusion Neural Networks. By providing both real and synthetic datasets, this task enables participants to analyze and compare the characteristics of synthetic images with their real counterparts, investigating potential "fingerprints" and patterns related to the training process.
More information will be added soon.
Evaluation Methodology
Information will be added soon.
Participant registration
Please refer to the general ImageCLEF registration instructions
Results
CEUR Working Notes
Paper Submission Instructions
The full schedule is available at:
https://clef2026.clef-initiative.eu/dates/
Important Dates
- 28.05.2026: Submission of participant papers [CEUR-WS]
- 30.06.2026: Notification of acceptance
- 21.09.2026: CLEF 2026, Jena, Germany
All submissions, reviews, and camera-ready versions will be handled through EasyChair: EasyChair CLEF 2026
A separate EasyChair track will be created for each lab/workshop. Please make sure you submit in the correct track.
The papers will go through a review process and will receive a decision from the lab organizers.
The participant papers should be written using the template provided here:
CLEF 2026 Working Notes Submission Template
Submissions are expected to be in English language and 5 pages minimum, with no maximum page limit.
Citations
When referring to ImageCLEFmedical GANs 2026, please cite the following:
@inproceedings{ImageCLEFmedicalGANs2026,
title = {Overview of the 2026 {I}mage{CLEF}medical {GAN}s task: Fingerprint detection, latent space analysis, and privacy-preserving medical image synthesis},
author = {Alexandra{-}Georgiana Andrei and Mihai{-}Gabriel Constantin and Mihai Dogariu and Dzmitry Karpenka and Yuri Prokopchuk and Ahmedkhan Radzhabov and Dan{-}Cristian Stanciu and Liviu{-}Daniel \c{S}tefan and Vassili Kovalev and Henning M{\"u}ller and Bogdan Ionescu},
booktitle = {CLEF 2026 Working Notes},
series = {CEUR Workshop Proceedings},
year = {2026},
month = {September 21--24},
address = {Jena, Germany},
publisher = {CEUR-WS.org},
}
When referring to ImageCLEF 2026, please cite the following:
@inproceedings{ImageCLEF2026,
title = {Overview of ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security},
author = {Bogdan Ionescu and Henning M{\"u}ller and Dan{-}Cristian Stanciu and Andrei Radu and Radu{-}George Bolborici and Marian Negru and Alexandru{-}Florin Ene and Vlad{-}Mihai Vasilescu and Ana-Antonia Nicolae and Liviu{-}Daniel \c{S}tefan and Mihai{-}Gabriel Constantin and Mihai Dogariu and Alexandra{-}Georgiana Andrei and Hendrik Damm and Tabea M. G. Pakull and Asma {Ben Abacha} and Alba {Garc\'ia Seco de Herrera} and Christoph M. Friedrich and Raphael Br{\"u}ngel and Lea Reinartz and Henning Sch{\"a}fer and Cynthia Sabrina Schmidt and Benjamin Bracke and Praveen Nath and Bahad{\i}r Ery{\i}lmaz and Maja Hjuler and Diandra Fabre and Claire Lemaire and Benjamin Lecouteux and Didier Schwab and Dimitar Dimitrov and Ming Shan Hee and Momina Ahsan and Sarfraz Ahmad and Dimitrina Zlatkova and Georgi Pachov and Zhuohan Xie and Preslav Nakov and Ivan Koychev and Juampablo E. {Heras Rivera} and Daniel K. Low and Wen{-}wai Yim and Jacob Ruzevick and Dan Child and Mehmet Kurt and Zhaoyi Sun and Fei Xia and Meliha Yetisgen and Ahmedkhan Radzhabov and Yuri Prokopchuk and Vassili Kovalev and Dzmitry Karpenka and Steven A. Hicks and Sushant Gautam and Michael A. Riegler and Vajira Thambawita and P\r{a}l Halvorsen and Mohammad {El Sakka} and Josiane Mothe and Alexandra B\u{a}icoianu and Corneliu{-}Nicolae Florea and Mihai Ivanovici},
booktitle = {Experimental IR Meets Multilinguality, Multimodality, and Interaction},
series = {Proceedings of the Seventeenth International Conference of the CLEF Association (CLEF 2026)},
year = {2026},
month = {September 21--24},
address = {Jena, Germany},
publisher = {Springer Lecture Notes in Computer Science LNCS},
}
Contact
Organizers:
- Alexandra Andrei, alexandra.andrei(at)upb.ro, National University of Science and Technology POLITEHNICA Bucharest, Romania
- Ahmedkhan Radzhabov, National Academy of Science of Belarus, Minsk, Belarus
- Yuri Prokopchuk, National Academy of Science of Belarus, Minsk, Belarus
- Liviu-Daniel Ștefan, National University of Science and Technology POLITEHNICA Bucharest, Romania
- Dzmitry Karpenka, National Academy of Science of Belarus, Minsk, Belarus
- Mihai Gabriel Constantin, National University of Science and Technology POLITEHNICA Bucharest, Romania.
- Dan-Cristian Stanciu, National University of Science and Technology POLITEHNICA Bucharest
- Mihai Dogariu, National University of Science and Technology POLITEHNICA Bucharest, Romania.
- Vassili Kovalev, vassili.kovalev(at)gmail.com, National Academy of Science of Belarus, Minsk, Belarus
- Bogdan Ionescu , bogdan.ionescu(at)upb.ro, National University of Science and Technology POLITEHNICA Bucharest, Romania
- Henning Müller , henning.mueller(at)hevs.ch, University of Applied Sciences Western Switzerland, Sierre, Switzerland
Acknowledgments