Welcome to the 1st edition of the CvTR QA Task!
Motivation
Charts and tables are among the most common ways of presenting structured information in scientific articles, reports, news and everyday documents. Understanding them requires more than recognizing text or visual elements: a system must read values, relate rows, columns and data series to each other, and reason over them to reach an answer. While visual question answering and table question answering have both received considerable attention, there is still a gap between these standard settings and real-world multilingual document understanding, where the same information may appear in different languages and in different forms, such as an image, source code or plain text.
The Multilingual Multimodal Chart and Tabular Question Answering (CvTR QA) task evaluates whether models can accurately extract, understand, and reason over structured information presented in charts and tables across languages and representations.
Preliminary Schedule
- TBD: Registration opens for all ImageCLEF tasks
- TBD: Development dataset released
- TBD: Test dataset released
- TBD: Registration closes for all ImageCLEF tasks
- TBD: Deadline for submitting participant runs
- TBD: Release of the processed results by the task organizers
- TBD: Submission of participant papers [CEUR-WS]
- TBD: Notification of acceptance
- TBD: CLEF 2027, TBD
Task Description
In 2027, ImageCLEF introduces the CvTR QA task, which targets question answering over charts and tables presented in multiple languages and in both visual and textual form. The questions cover numerical reading, trend identification, category comparison, field matching, cross-row and cross-column association, conditional filtering, and multi-step reasoning. The task is split into four subtasks:
Subtask 1: Chart Image QA
Participants will answer questions using chart images as input. The subtask covers around 40 chart types, including bar charts, line charts, pie charts, scatter plots, composite charts, and infographics.
Subtask 2: Chart Code QA
Participants will answer questions using chart source code as textual input. The subtask evaluates the understanding of chart structure, data mappings, and visual semantics.
Subtask 3: Visual Table QA
Participants will answer questions using table images as input. The subtask focuses on layout understanding, row-column relations, merged cells, and hierarchical headers.
Subtask 4: Textual Table QA
Participants will answer questions using plain-text tables in formats such as HTML, LaTeX, Markdown, and CSV. The subtask focuses on structured information extraction and reasoning.
Data
The task combines visual and textual or code-based representations of charts and tables. The current version of the dataset contains approximately 2,000 question-answer pairs across around six languages, including Chinese, English, Arabic, and Bulgarian.
More information will be added soon.
Evaluation Methodology
The task provides a unified evaluation of cross-lingual reasoning, multimodal parsing, and robustness across equivalent visual and textual representations of the same information.
More information will be added soon.
Participant registration
Please refer to the general ImageCLEF registration instructions
Results
CEUR Working Notes
Citations
Contact
Organizers:
- Zhuohan Xie, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE
- Preslav Nakov, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE
- Momina Ahsan, Mohamed bin Zayed University of Artificial Intelligence (MBZUAI), UAE
- Dimitar Dimitrov, Sofia University "St. Kliment Ohridski", Bulgaria