Welcome to the 10th edition of the Caption Task!
Motivation
Interpreting and summarizing the insights gained from medical images such as radiology output is a time-consuming task that involves highly trained experts and often represents a bottleneck in clinical diagnosis pipelines.
Consequently, there is a considerable need for automatic methods that can approximate this mapping from visual information to condensed textual descriptions. The more image characteristics are known, the more structured are the radiology scans and hence, the more efficient are the radiologists regarding interpretation. We work on the basis of a large-scale collection of figures from open access biomedical journal articles (PubMed Central). All images in the training data are accompanied by UMLS concepts extracted from the original image caption.
Lessons learned:
- In the first and second editions of this task, held at ImageCLEF 2017 and ImageCLEF 2018, participants noted a broad variety of content and situation among training images. In 2019, the training data was reduced solely to radiology images, with ImageCLEF 2020 adding additional imaging modality information, for pre-processing purposes and multi-modal approaches.
- The focus in ImageCLEF 2021 lay in using real radiology images annotated by medical doctors. This step aimed at increasing the medical context relevance of the UMLS concepts, but more images of such high quality are difficult to acquire.
- As uncertainty regarding additional source was noted, we will clearly separate systems using exclusively the official training data from those that incorporate additional sources of evidence.
- For ImageCLEF 2022, an extended version of the ImageCLEF 2020 dataset was used. For the caption prediction subtask, a number of different additional evaluations metrics were introduced with the goal of replacing the primary evaluation metric in future iterations of the task.
- For ImageCLEF 2023, several issues with the dataset (large number of concepts, lemmatization errors, duplicate captions) were tackled and based on experiments in the previous year, BERTScore was used as the primary evaluation metric for the caption prediction subtask.
Schedule
- 01.12.2025: Website goes live
- 26.01.2026: Registration opens
- 15.02.2026: Development dataset released
- 19.04.2026: Test dataset released
- 30.04.2026: Deadline for submitting participant runs
- 07.05.2026: Release of processed results by the task organizers
- 28.05.2026: Submission of participant papers (CEUR-WS)
- 30.06.2026: Notification of acceptance
- 06.07.2026: Camera Ready Submission Deadline
Task Description
For captioning, participants will be requested to develop solutions for automatically identifying individual components from which captions are composed in Radiology Objects in COntext version 2[2] images. ImageCLEFmedical Caption 2026 consists of the following subtasks:
Standard Tasks
Synthetical Tasks NEW
New this year, we introduce two synthetical subtasks. For these tasks, we generated new captions and derived new concepts from those captions. The evaluation methodology remains the same as for the standard tasks, with the exception that there is no secondary score for the Synthetical Concept Detection subtask.
Important: Datasets must not be mixed between the standard and synthetical tasks.
Explainability Task
In addition, we ask participants to provide explanations for the captions of a small subset (will be released with the test dataset) of images. We encourage people to be creative. There are no technical limitations to this task. The explanations will be manually evaluated by a radiologist for interpretability, relevance, and creativity. Examples of how such an explanation might look like are provided as follows:

Submission Limits
- Development set: Up to 10 submissions (maximum 5 per day)
- Test set: Up to 3 submissions
Baselines
We provide trained baselines for reference:
- Caption Prediction: Qwen3-4B-VL, fine-tuned with Unsloth
- Concept Detection: EfficientNet-based model
Concept Detection Task
The first step to automatic image captioning and scene understanding is identifying the presence and location of relevant concepts in a large corpus of medical images. Based on the visual image content, this subtask provides the building blocks for the scene understanding step by identifying the individual components from which captions are composed. The concepts can be further applied for context-based image and information retrieval purposes.
Evaluation is conducted in terms of set coverage metrics such as precision, recall, and combinations thereof.
Caption Prediction Task
On the basis of the concept vocabulary detected in the first subtask as well as the visual information of their interaction in the image, participating systems are tasked with composing coherent captions for the entirety of an image. In this step, rather than the mere coverage of visual concepts, detecting the interplay of visible elements is crucial for strong performance.
This year, we will use BERTScore as the primary evaluation metric and ROUGE as the secondary evaluation metric for the caption prediction subtask. Other metrics such as MedBERTScore, MedBLEURT, and BLEU will also be published.
Data
The data for the caption task will contain curated images from the medical literature including their captions and associated UMLS terms that are manually controlled as metadata. A more diverse data set will be made available to foster more complex approaches.
For questions regarding the dataset please use the challenge website forum or contact hendrik.damm@fh-dortmund.de.
For the development dataset, Radiology Objects in COntext Version 2 (ROCOv2) [2], an updated and extended version of the Radiology Objects in COntext (ROCO) dataset [1], is used for both subtasks. As in previous editions, the dataset originates from biomedical articles of the PMC OpenAccess subset, with the test set comprising a previously unseen set of images.
- Training Set: Consists of tba radiology images
- Validation Set: Consists of tba radiology images
- Test Set: Consists of tba radiology images
Concept Detection Task
The concepts were generated using a reduced subset of the UMLS 2022 AB release. To improve the feasibility of recognizing concepts from the images, concepts were filtered based on their semantic type. Concepts with low frequency were also removed, based on suggestions from previous years.
Caption Prediction Task
For this task each caption is pre-processed in the following way:
- removal of links from the captions
Evaluation methodology
The source code of the evaluation is available on Github (https://github.com/taubsity/clef-caption-evaluation-2026).
The repository also provides scripts to check submission formats.
For questions regarding the evaluation scripts please use the challenge website forum or contact tabea.pakull@uk-essen.de.
Concept Detection
Evaluation is conducted in terms of F1 scores between system predicted and ground truth concepts, using the following methodology and parameters:
- The default implementation of the Python scikit-learn (v0.17.1-2) F1 scoring method is used. It is documented here.
- A Python (3.x) script loads the candidate run file, as well as the ground truth (GT) file, and processes each candidate-GT concept sets.
- For each candidate-GT concept set, the y_pred and y_true arrays are generated. They are binary arrays indicating for each concept contained in both candidate and GT set if it is present (1) or not (0).
- The F1 score is then calculated. The default 'binary' averaging method is used.
- All F1 scores are summed and averaged over the number of elements in the test set, giving the final score.
- The primary score considers any concept. The secondary score filters both predicted and GT concepts to the set of manually annotated concepts before repeating the same F1 scoring steps. Note: There is no secondary score for the Synthetical Concept Detection subtask.
The ground truth for the test set was generated based on the same reduced subset of the UMLS 2022 AB release which was used for the training data (see above for more details).
Caption Prediction
This year, ranking of participants is based on an average score over all used metrics. In total 6 metrics are computed that fall into the aspects of relevance and factuality.
Relevance
In order to evaluate the relevance aspect of generated captions the following metrics are used:
- Image and Caption Similarity
- BERT-Score (Recall) with inverse document frequency (idf) scores computed from the test corpus for importance weighting
- Recall-Oriented Understudy for Gisting Evaluation (ROUGE) for overlap of unigrams (ROUGE-1) (F-measure)
- Bilingual Evaluation Understudy with Representations from Transformers (BLEURT)
Image and Caption Similarity is computed using the following methodology:
Using a medical imaging embedding model for calculating embeddings of the caption and the image and calculating similarity of these embeddings.
Note: For the following relevance metrics (BERT-Score, ROUGE and BLEURT), each caption is pre-processed in the same way:
- The caption is converted to lower-case.
- Replace numbers with the token 'number'.
- Remove punctuation.
Note that the captions are always considered as a single sentence, even if it actually contains several sentences.
BERTScore is calculated using the following methodology and parameters:
The native Python implementation of BERTScore is used. This scoring method is based on the paper "BERTScore: Evaluating Text Generation with BERT" and aims to measure the quality of generated text by comparing it to a reference. We use Recall BERTScore with inverse document frequency (idf) scores computed from the test corpus for importance weighting as this setting correlates the most with human ratings for the image captioning task reported in the BERTScore paper.
To calculate BERTScore, we use the microsoft/deberta-xlarge-mnli model, which can be found on the Hugging Face Model Hub. The model is pretrained on a large corpus of text and fine-tuned for natural language inference tasks. It can be used to compute contextualized word embeddings, which are essential for BERTScore calculation.
To compute the final BERTScore, we first calculate the individual score (Recall idf) for each caption. The BERTScore is then averaged across all captions to give the final score.
The ROUGE score is calculated using the following methodology and parameters:
The native python implementation of ROUGE scoring method is used. It is designed to replicate results from the original perl package that was introduced in the paper "ROUGE: A Package for Automatic Evaluation of Summaries".
Specifically, we calculate the ROUGE-1 (F-measure) score, which measures the number of matching unigrams between the model-generated text and a reference. The final score is the average ROUGE-1 over all captions.
For calculation of BLEURT the following methodology and parameters are used:
The native Python implementation of BLEURT is used. This scoring method is based on the paper "BLEURT: Learning Robust Metrics for Text Generation". The aim of BLEURT is to provide an evaluation metric for text generation by learning from human judgments using BERT-based representations. In this evaluation, the recommended BLEURT-20 checkpoint is employed.
Factuality
In order to evaluate the factuality aspect of generated captions the following metrics are used:
- Unified Medical Language System (UMLS) Concept F1
- AlignScore
We calculate the UMLS F1 using the following methodology:
We use MedCAT to get the medical entities (UMLS concepts) of the caption and the predicted caption. We only match entities with the semantic types that are used to calculate MEDCON as described in the paper "Aci-bench: a Novel Ambient Clinical Intelligence Dataset for Benchmarking Automatic Visit Note Generation".
The AlignScore is calculated using the following methodology and parameters:
The native Python implementation of the AlignScore is used. It implements the metric based on RoBERTa introduced in the paper "AlignScore: Evaluating Factual Consistency with A Unified Alignment Function". The checkpoints are available on the Huggingface Model Hub. The AlignScore is designed to evaluate factual consistency in text generation by assessing the alignment of information between two pieces of text. The model calculates a score by splitting long contexts into manageable chunks and matching each claim sentence with the most supportive context chunk. The final AlignScore is the average alignment score across all claim sentences.
Certificates
This year, we will provide certificates for participating teams that submit working notes, as well as for the best-ranked teams per subtask.
Participant registration
Please refer to the general ImageCLEF registration instructions
Results
Note to the participants: Use the official leaderboards on sciebo to report your results.
Explainability:
The complete evaluation files are available here: https://portal.uk-essen.de/enqsig/link?id=BAgAAAAJv8BhoGWDRokAAADJpWasQj89hCJV7oZDXFby4doGcgUknLjkoPsElm0v-xlwVtTUHVlucxwcIY7_4tCPqGTp8WDQSM1Ji2tyvQnJ71xtt1mMEdp_fMM_mDFkIkSVWfdqxJfEHNhMOFzbH9rIhsRAW5rUhAZptb1gCWWfvTKBI5KfCGuoCQ7aPdFyj9zXJ46spqa6Jw2
Password: ALMZHQd5sD
Caption, Caption Synthetic, Concept, Concept Synthetic:
The complete evaluation files are available here: https://portal.uk-essen.de/enqsig/link?id=BAgAAACWnrM8NrmR24kAAABC9CxGTgB9jx8FdIN5HLdRQgCwyFE-u25ycZSM49XTQ3WX_7-xn5rmUlzcAIupiz23VheFZXHGqnlxa7HNaaFp86_QpOlljytIGkFNXJOmYdsOfd2n9_-zPNd-jAHHdKk6wpyzvW_UKyqZpwxLMXrXYPwtEOWSkt05jZIkC6LKGIiSiXfkSwb4ig2
Password: clef_caption_2026
Concept Detection
| Rank |
Owner |
Score |
Score_secondary |
| 1 |
dsgt_caption |
0.5790 |
0.9657 |
| 2 |
archimedes |
0.5781 |
0.9591 |
| 3 |
UIT_Worker |
0.5725 |
0.9533 |
| 4 |
basharderar |
0.5725 |
0.9419 |
| 5 |
AUEB/DMST |
0.5721 |
0.9503 |
| 6 |
nhanbui |
0.5713 |
0.9553 |
| 7 |
DeepLens |
0.5648 |
0.9336 |
| 8 |
dingglebellw |
0.5606 |
0.9428 |
| 9 |
NUSTPB |
0.5594 |
0.9465 |
| 10 |
krishnatewari |
0.5234 |
0.8403 |
| 11 |
UIT_HighDefinition |
0.2739 |
0.9349 |
Concept Detection (Synthetic)
| Rank |
Owner |
Score |
Score_secondary |
| 1 |
AUEB/DMST |
0.5609 |
0 |
| 2 |
krishnatewari |
0.4741 |
0 |
| 3 |
NUSTPB |
0.1048 |
0 |
Caption Prediction
| Rank |
Owner |
Overall |
Relevance |
Factuality |
Bert |
Rouge |
Similarity |
Bleurt |
Medcat |
Align |
| 1 |
AIstatLab |
0.3755 |
0.5786 |
0.1724 |
0.6250 |
0.3234 |
1.0104 |
0.3555 |
0.2160 |
0.1287 |
| 2 |
UFPE-Essex-UNED |
0.3603 |
0.5527 |
0.1679 |
0.6017 |
0.2725 |
0.9902 |
0.3463 |
0.1822 |
0.1537 |
| 3 |
dsgt_caption |
0.3571 |
0.5365 |
0.1777 |
0.6042 |
0.2809 |
0.9381 |
0.3229 |
0.1943 |
0.1610 |
| 4 |
AUEB/DMST |
0.3516 |
0.5288 |
0.1743 |
0.6130 |
0.2728 |
0.9113 |
0.3183 |
0.1812 |
0.1674 |
| 5 |
csmorgan |
0.3458 |
0.5377 |
0.1539 |
0.6075 |
0.2725 |
0.9408 |
0.3299 |
0.1751 |
0.1327 |
| 6 |
gfreitas |
0.3343 |
0.5201 |
0.1484 |
0.6013 |
0.2226 |
0.9273 |
0.3293 |
0.1563 |
0.1406 |
| 7 |
UIT_HighDefinition |
0.3193 |
0.4999 |
0.1388 |
0.5977 |
0.2243 |
0.8496 |
0.3278 |
0.1397 |
0.1379 |
| 8 |
NUSTPB |
0.3065 |
0.4705 |
0.1426 |
0.5465 |
0.2224 |
0.7787 |
0.3342 |
0.1443 |
0.1408 |
| 9 |
UMUTeam |
0.2990 |
0.4798 |
0.1181 |
0.5670 |
0.2230 |
0.8147 |
0.3146 |
0.1048 |
0.1315 |
| 10 |
DS-MED Lab |
0.2814 |
0.4526 |
0.1102 |
0.5690 |
0.2126 |
0.7286 |
0.3004 |
0.1427 |
0.0778 |
| 11 |
krishnatewari |
0.2685 |
0.4346 |
0.1025 |
0.5398 |
0.2059 |
0.6788 |
0.3137 |
0.1219 |
0.0831 |
| 12 |
UACH-VisionLab |
0.2673 |
0.4275 |
0.1072 |
0.5742 |
0.1818 |
0.6589 |
0.2951 |
0.1155 |
0.0989 |
| 13 |
nguyenthinh |
0.2527 |
0.4373 |
0.0680 |
0.5720 |
0.2207 |
0.6514 |
0.3051 |
0.0872 |
0.0488 |
Caption Prediction (Synthetic)
| Rank |
Owner |
Overall |
Relevance |
Factuality |
Bert |
Rouge |
Similarity |
Bleurt |
Medcat |
Align |
| 1 |
AIstatLab |
0.5840 |
0.7183 |
0.4496 |
0.7580 |
0.6353 |
0.9737 |
0.5063 |
0.5332 |
0.3660 |
| 2 |
AUEB/DMST |
0.5136 |
0.6542 |
0.3731 |
0.7195 |
0.5555 |
0.8902 |
0.4515 |
0.4677 |
0.2784 |
| 3 |
UMUTeam |
0.4632 |
0.6148 |
0.3116 |
0.6930 |
0.5127 |
0.8159 |
0.4376 |
0.4099 |
0.2133 |
| 4 |
UFPE-Essex-UNED |
0.4625 |
0.6372 |
0.2878 |
0.6621 |
0.4795 |
0.9843 |
0.4230 |
0.3782 |
0.1974 |
| 5 |
krishnatewari |
0.4341 |
0.5758 |
0.2924 |
0.6687 |
0.4758 |
0.7412 |
0.4172 |
0.3775 |
0.2072 |
| 6 |
madhab |
0.3861 |
0.5200 |
0.2522 |
0.6033 |
0.3889 |
0.7406 |
0.3473 |
0.3412 |
0.1633 |
| 7 |
NUSTPB |
0.3095 |
0.4786 |
0.1404 |
0.5247 |
0.2670 |
0.7838 |
0.3388 |
0.1264 |
0.1543 |
Explainability Task
| Rank |
Team |
Readability |
Accuracy |
Level of detail |
Focus - Caption |
Consistency |
Comprehens. |
Focus - Visual. |
Method |
Clinician’s Favorite |
Overall |
| 1 |
Mahmudul Hoque, Morgan State University |
4.6 |
3.6 |
3.4 |
4.3 |
3.1 |
2.9 |
3.4 |
3 |
5 |
3.69 |
| 2 |
Krishna Tewari, Indian Institute of Technology |
2.625 |
2.125 |
2.75 |
2.625 |
2.625 |
3.125 |
2.75 |
4 |
3 |
2.85 |
| 3 |
Gabriel Freitas, Universidade Estadual de Campinas |
4.5 |
3.1 |
3.6 |
4.1 |
1.0 |
2.1 |
1.0 |
3 |
3 |
2.83 |
CEUR Working Notes
The working-notes paper is your opportunity to describe your approach, present all submitted runs and discuss the results. All participating teams with at least one graded submission, regardless of the score, should submit a CEUR working notes paper. Teams who participated in both tasks should generally submit only one report.
Use the official CEUR-WS template
https://clef2026.clef-initiative.eu/calls/submitting/#labs-working-notes--task-overview-papers
Make sure the EasyChair metadata (author names and order, title, affiliations) exactly match the PDF, as these fields feed directly into the proceedings.
- Camera-ready + signed copyright form:28.05.2026
Dataset image attribution
If you include dataset images in your paper, you must provide the correct attribution. Use the lookup file below to find the attribution string for each image ID:
https://fh-dortmund.sciebo.de/s/KT9AMjPtoq3pxTz
Insert the attribution in the figure caption or directly beside the image.
Working Notes Papers should cite both the ImageCLEF 2026 overview paper as well as the ImageCLEFmedical task overview paper and the ROCOv2 dataset paper, citation information is available in the Citations section below.
Reproducibility
We encourage you to make your work as reproducible as possible by releasing code, trained models, and detailed instructions on a public repository (e.g. GitHub) and pointing to it in your paper.
ArXiv references
Please refrain from citing preprints (e.g. arXiv) without checking if they have been published in the meantime. The publication details proof the value of the cited work and give attribution to the authors; an arXiv reference is only a preprint with no peer review. You can try https://preprintresolver.eu/ to give proper attribution to authors and improve the quality of your work. If no proper reference/doi is found, you can use the arXiv reference.
Citations
It is mandatory to cite the overview papers and the dataset.
ImageCLEFmedical Overview 2026
Hendrik Damm et al. "Overview of ImageCLEFmedical 2026 -- Medical Concept Detection and Caption Generation with Synthetic Data Extensions."
@inproceedings{ImageCLEFmedicalCaptionOverview2026,
author = {Damm, Hendrik and Pakull, Tabea M. G. and Reinartz, Lea and Bracke, Benjamin and Ery{\i}lmaz, Bahad{\i}r and Nath, Praveen and Br{\"u}ngel, Raphael and Schmidt, Cynthia S. and Sch{\"a}fer, Henning and Ben Abacha, Asma and {Garc{\'\i}a Seco de Herrera}, Alba and M{\"u}ller, Henning and Friedrich, Christoph M.},
title = {Overview of {ImageCLEFmedical} 2026 -- Medical Concept Detection and Caption Generation with Synthetic Data Extensions},
booktitle = {CLEF2026 Working Notes},
series = {{CEUR} Workshop Proceedings},
year = {2026},
publisher = {CEUR-WS.org},
month = {September 21-24},
address = {Jena, Germany},
note = {To appear}
}
ImageCLEF 2026 General Overview
Bogdan Ionescu et al. "Overview of ImageCLEF 2026: Multimedia Retrieval in Medical, Social Media and Content Recommendation Applications."
@inproceedings{ImageCLEF2026,
title = {Overview of ImageCLEF 2026: Multimodal Challenges in Medicine, Science, Agritech, and Security},
author = {Bogdan Ionescu and Henning M{\"u}ller and Dan{-}Cristian Stanciu and Andrei Radu and Radu{-}George Bolborici and Marian Negru and Alexandru{-}Florin Ene and Vlad{-}Mihai Vasilescu and Ana-Antonia Nicolae and Liviu{-}Daniel \c{S}tefan and Mihai{-}Gabriel Constantin and Mihai Dogariu and Alexandra{-}Georgiana Andrei and Hendrik Damm and Tabea M. G. Pakull and Asma {Ben Abacha} and Alba {Garc\'ia Seco de Herrera} and Christoph M. Friedrich and Raphael Br{\"u}ngel and Lea Reinartz and Henning Sch{\"a}fer and Cynthia Sabrina Schmidt and Benjamin Bracke and Praveen Nath and Bahad{\i}r Ery{\i}lmaz and Maja Hjuler and Diandra Fabre and Claire Lemaire and Benjamin Lecouteux and Didier Schwab and Dimitar Dimitrov and Ming Shan Hee and Momina Ahsan and Sarfraz Ahmad and Dimitrina Zlatkova and Georgi Pachov and Zhuohan Xie and Preslav Nakov and Ivan Koychev and Juampablo E. {Heras Rivera} and Daniel K. Low and Wen{-}wai Yim and Jacob Ruzevick and Dan Child and Mehmet Kurt and Zhaoyi Sun and Fei Xia and Meliha Yetisgen and Ahmedkhan Radzhabov and Yuri Prokopchuk and Vassili Kovalev and Dzmitry Karpenka and Steven A. Hicks and Sushant Gautam and Michael A. Riegler and Vajira Thambawita and P\r{a}l Halvorsen and Mohammad {El Sakka} and Josiane Mothe and Alexandra B\u{a}icoianu and Corneliu{-}Nicolae Florea and Mihai Ivanovici},
booktitle = {Experimental IR Meets Multilinguality, Multimodality, and Interaction},
series = {Proceedings of the Seventeenth International Conference of the CLEF Association (CLEF 2026)},
year = {2026},
month = {September 21--24},
address = {Jena, Germany},
publisher = {Springer Lecture Notes in Computer Science LNCS},
}
}
ROCOv2 Dataset
Johannes Rückert et al. "ROCOv2: Radiology Objects in Context Version 2, an Updated Multimodal Image Dataset."
@article{2405.10004v2,
title = {{ROCOv2}: Radiology Objects in COntext Version 2, an Updated Multimodal Image Dataset},
author = {Johannes R{\"u}ckert and Louise Bloch and Raphael Br{\"u}ngel and Ahmad Idrissi{-}Yaghir and Henning Sch{\"a}fer and Cynthia S. Schmidt and Sven Koitka and Obioma Pelka and Asma Ben Abacha and Alba Garc{'{\i}}a Seco de Herrera and Henning M{\"u}ller and Peter Horn and Felix Nensa and Christoph M. Friedrich},
journal = {Scientific Data},
volume = {11},
number = {1},
year = {2024},
doi = {10.1038/s41597-024-03496-6}
}
Contact
Organizers:
- Hendrik Damm <hendrik.damm(at)fh-dortmund.de>, University of Applied Sciences and Arts Dortmund, Germany
- Tabea M. G. Pakull, <tabea.pakull(at)uk-essen.de>, Institute for Transfusion Medicine, University Hospital Essen, Germany
- Asma Ben Abacha <abenabacha(at)microsoft.com>, Microsoft, USA
- Alba García Seco de Herrera <alba.garcia(at)essex.ac.uk>, University of Essex, UK
- Christoph M. Friedrich <christoph.friedrich(at)fh-dortmund.de>, University of Applied Sciences and Arts Dortmund, Germany
- Henning Müller <henning.mueller(at)hevs.ch>, University of Applied Sciences Western Switzerland, Sierre, Switzerland
- Raphael Brüngel <raphael.bruengel(at)fh-dortmund.de>, University of Applied Sciences and Arts Dortmund, Germany
- Lea Reinartz, University of Applied Sciences and Arts Dortmund, Germany
- Praveen Nath, University of Applied Sciences and Arts Dortmund, Germany
- Henning Schäfer <henning.schaefer(at)uk-essen.de>, Institute for Transfusion Medicine, University Hospital Essen, Germany
- Cynthia S. Schmidt, Institute for Artificial Intelligence in Medicine (IKIM), University Hospital Essen
- Benjamin Bracke, University of Applied Sciences and Arts Dortmund, Germany
- Bahadir Eryilmaz, Institute for Artificial Intelligence in Medicine, Germany
Acknowledgments
[1] Rückert, J., Bloch, L., Brüngel, R., Idrissi-Yaghir, A., Schäfer, H., Schmidt, C. S., Koitka, S., Pelka, O., Abacha, A. B., de Herrera, A. G. S., Müller, H., Horn, P. A., Nensa, F., & Friedrich, C. M. (2024). ROCOv2: Radiology objects in COntext version 2, an updated multimodal image dataset. https://doi.org/10.48550/ARXIV.2405.10004