You are here

ImageCLEFmed GANs

Welcome to the 5th edition of the GANs task!

Motivation

Description

AI and computer-aided diagnosis systems for medical tasks such as prediction, detection, and classification depend on access to large and diverse datasets for effective training. High-quality data enable these models to learn complex patterns, improving both accuracy and reliability. However, obtaining real medical data is difficult due to strict privacy regulations and ethical concerns. Patients typically consent to the use of their medical information for their own clinical care, but not for broad research purposes. As a result, collecting sufficiently large datasets for training AI models remains a significant challenge and slows progress in developing advanced healthcare tools.

Synthetic medical data has emerged as a promising solution. Generative models such as GANs (Generative Adversarial Networks) can produce realistic-looking data that resemble true medical images while not being tied to any specific patient. These synthetic datasets allow researchers to build and evaluate AI systems without directly relying on sensitive patient information, potentially easing data-access barriers and enabling greater diversity and scalability.

However, an important concern remains: synthetic data must be free of hidden traces or “fingerprints” from the real images used during training. If synthetic images inadvertently reproduce patient-specific details, they could still leak private information. Ensuring that generative models do not memorize or reveal identifiable patterns is therefore crucial. Only truly privacy-preserving synthetic data can safely accelerate innovation in AI-powered healthcare.

Lessons learned

  • In the first three editions of this task (2023-2025), various generative models were analyzed within the framework of the first subtask to investigate whether synthetic images contained "fingerprints" of the real medical data used during training. The results demonstrated that generative models do retain and imprint features from their training data, raising important security and privacy concerns. These findings underscore the need for robust techniques to detect and mitigate such imprints to ensure that synthetic images protect patient privacy while maintaining their utility for research and development.

  • In the 2nd edition of the task (2024), it was confirmed that generative models leave unique "fingerprints" on the synthetic images they produce. By analyzing images generated from various models, distinct patterns and features were identified that allowed the attribution of synthetic images to their respective generative models.

  • In the 3rd edition (2025), subtask “Identify Training Data Subsets” explored attribution at a broader scale: identifying the training subset used to generate synthetic images from a diffusion model. Identification performance was significantly higher, with multiple teams exceeding 98% accuracy. While this task did not involve image-level membership inference, it revealed that diffusion-generated images can still reflect statistical characteristics of the training data subsets. The success of various supervised and semi-supervised classification techniques suggests that diffusion models, while effective at preserving image realism, may still encode latent information about their source data distributions. This raises important considerations about the use of synthetic data for data sharing or augmentation, particularly when the provenance or diversity of training data must remain confidential.

  • In the 4th edition (2026), subtask “Detect training data usage” remained highly challenging, with all submissions performing close to chance level. In contrast, subtask “Identify corresponding latent vectors” was solved with up to 100% accuracy, showing that the mapping between latent vectors and generated images can be highly predictable and therefore a potential source of privacy leakage. Subtask “Privacy-preserving CT slice generation” revealed a clear trade-off between privacy preservation and image realism: the submission with the highest privacy score produced images that diverged substantially from the real CT distribution, while submissions with more realistic images obtained lower privacy scores.

News

Preliminary Schedule

    • TBD: Registration opens for all ImageCLEF tasks
    • TBD: Development dataset released (depends on task)
    • TBD: Test dataset released (depends on task)
    • TBD: Registration closes for all ImageCLEF tasks
    • TBD: Deadline for submitting participant runs
    • TBD: Release of the processed results by the task organizers (depends on task)
    • TBD: Submission of participant papers [CEUR-WS]
    • TBD: Notification of acceptance
    • TBD: CLEF 2027, TBD

Task Description

In 2027, we continue the traditional ImageCLEFmed GAN challenge task on detecting whether real medical images were used to train generative models. Together, these subtasks aim to deepen our understanding of how generative models memorize, reconstruct, or unintentionally expose sensitive biomedical information.

Subtask 1: TBD

Subtask 2: TBD

Subtask 3: TBD

Data

Information will be added soon.

Evaluation Methodology

Information will be added soon.

Participant registration

Please refer to the general ImageCLEF registration instructions

Results

CEUR Working Notes

Citations

Contact

Organizers:

  • Alexandra Andrei, alexandra.andrei(at)upb.ro, National University of Science and Technology POLITEHNICA Bucharest, Romania
  • Ahmedkhan Radzhabov, National Academy of Science of Belarus, Minsk, Belarus
  • Yuri Prokopchuk, National Academy of Science of Belarus, Minsk, Belarus
  • Liviu-Daniel Ștefan, National University of Science and Technology POLITEHNICA Bucharest, Romania
  • Dzmitry Karpenka, National Academy of Science of Belarus, Minsk, Belarus
  • Mihai Gabriel Constantin, National University of Science and Technology POLITEHNICA Bucharest, Romania.
  • Dan-Cristian Stanciu, National University of Science and Technology POLITEHNICA Bucharest
  • Mihai Dogariu, National University of Science and Technology POLITEHNICA Bucharest, Romania.
  • Vassili Kovalev, vassili.kovalev(at)gmail.com, National Academy of Science of Belarus, Minsk, Belarus
  • Bogdan Ionescu , bogdan.ionescu(at)upb.ro, National University of Science and Technology POLITEHNICA Bucharest, Romania
  • Henning Müller , henning.mueller(at)hevs.ch, University of Applied Sciences Western Switzerland, Sierre, Switzerland

Acknowledgments