MAIBAI Stakeholder Workshop Explores Pathways Toward Trustworthy Medical AI
Thursday, 16th July 2026, Physikalisch-Technische Bundesanstalt, Berlin, (Germany), (Hybrid)
The third stakeholder workshop of the EMP Project 22HLT05 MAIBAI brings together experts from metrology, medicine, and AI to discuss the main results achieved by project, providing a comprehensive evaluation framework for AI in diagnostic imaging, using breast cancer screening as an exemplar. Every stage required to evaluate AI tools in mammography imaging is considered, spanning data sub-categorisation, augmentation of training and testing datasets with synthetic data, and assessment, validation and explainability of AI model outputs in clinical context.
The workshop is articulated in four thematic sessions:
An Introduction to MAIBAI, Mammography Data and Evaluation Platform
Brief introduction to the workshop – An overview of the main results achieved by the MAIBAI project
Alessandra Manzin, Istituto Nazionale di Ricerca Metrologica (INRiM), Torino, Italy
This workshop aims at discussing the main results achieved by the European Metrology Partnership Project 22HLT05 MAIBAI, providing a comprehensive evaluation framework for AI in diagnostic imaging, using breast cancer screening as an exemplar. Every stage required to evaluate AI tools in mammography imaging is considered, spanning data sub-categorisation, augmentation of training and testing datasets with synthetic data, and assessment, validation and explainability of AI model outputs in clinical context.
MAIBAI AI evaluation pipeline
Valentin Kraft, Fraunhofer Institute for Digital Medicine (MEVIS), Bremen, Germany
This talk deals with the design of an AI evaluation platform targeted for mammography. It will be demonstrated how an AI model — using the NYU breast cancer classification model as an example — can be evaluated using the MAIBAI evaluation pipeline on Google Cloud.
Subcategorisation of Data for AI Models in Mammography
Nadia Smith, Royal Surrey NHS Foundation Trust, Guildford, United Kingdom
AI reliability in medical imaging depends on capturing real-world variability, not just large datasets. This talk highlights key clinical and technical factors that influence mammography images and their potential impact on AI performance. We show how data subcategorisation improves traceability, validation, and equity, and outline practical approaches for incorporating it into model development and evaluation. Using the OPTIMAM Image Database (OMI-DB) as a case study, we also review current data subgroups and gaps in coverage.
Factors Influencing the Performance of AI when Introduced in Breast Cancer Screening
Alistair Mackenzie, Royal Surrey NHS Foundation Trust, Guildford, United Kingdom
The performance of an artificial intelligence (AI) model for detection of breast cancer was evaluated deploying the algorithm as is, into different breast screening units. We have shown variation in the implementation between screening centres, equipment and over time. Differences in the training set from the local data can cause issues. AI model will need to be tested on local data and carefully validated, to ensure rare cancers and small populations are treated equitably. Overall, we have demonstrated the need for fine tuning and type testing based on the clinical environment and post-deployment monitoring.
Uncertainty Quantification and Generative AI for Mammography
Uncertainty Quantification for Generative AI Systems in Mammography
Spencer Thomas, National Physical Laboratory (NPL), Teddington, United Kingdom
We will discuss various methods for quantifying uncertainty in generative modelling with AI using mammography as a use case. We will summarise the pros and cons of the different methods and also discuss metrics for assessing reliability of uncertainties and identification of anomalies.
Exploitation of Generative AI in Breast Cancer Detection
Alessandra Manzin, Istituto Nazionale di Ricerca Metrologica (INRiM), Torino, Italy
One methodology to supplement clinical databases for improving representativeness and robustness of AI models for diagnostic imaging is the use of synthetic data. Their generation can address domain adaptation challenges due to the use of different manufacturer hardware, software, and scanner protocols across medical institutions, representing an instrument for bias mitigation and reliability enhancement.
One of the generative AI models investigated in the context of mammography is neural style transfer, used to balance training datasets in terms of features strongly dependent on vendor post-processing software. In this talk, we will demonstrate its utility in improving the accuracy and confidence of AI models for breast cancer
detection.
Quality and Explainability of AI Systems in Mammography
Quality assessment of AI systems in mammography screening
Danny Panknin, Physikalisch-Technische Bundesanstalt (PTB), Berlin, Germany
When AI systems for mammography screening move toward clinical deployment, rigorous quality testing is essential to guarantee ‘safe’ use. While standard diagnostic accuracy (sensitivity and specificity) is the baseline, evaluating an AI system must go much further to ensure it is safe and trustworthy. The system must remain robust and fair across different imaging hardware and diverse patient demographics, provide calibrated model uncertainty estimates, and, if included, its explanations of predictions (e.g., through pixel highlighting) must be verified.
An important aspect is to measure the quality of existing AI models across multiple quality dimensions by designing and conducting controlled, large-scale experiments. This talk will provide an overview of comprehensive model quality testing, explain how to align these tests with clinical boundary conditions, and demonstrate key insights from our evaluation. Furthermore, we will discuss the practical and theoretical hurdles encountered in this process and provide actionable guidance deduced from the lessons learned.
Assessing explanation correctness in mammography screening
Stefan Haufe, Physikalisch-Technische Bundesanstalt (PTB), Berlin, Germany
To build trust in AI-assisted radiology, modern algorithms often provide visual explanations—such as heatmaps highlighting specific regions on a mammogram—to justify their predictions. However, for these Explainable AI (XAI) tools to be clinically
valuable, their explanations must inform users about the relationship between individual image structures and the predicted outcome, as well as the resulting role of each within the AI model.
This talk will outline a comprehensive framework for validating the correctness of XAI. We will discuss the entire evaluation pipeline: from defining a clinically relevant XAI objective (e.g., a statistical association between highlighted structures in an
image and the presence of cancer), through assembling specialized test datasets with verified anatomical ground truth, to selecting evaluation metrics that accurately measure whether this objective is met.
Explainability of AI Systems in Mammography for Radiology Workflows
The Plausibility Trap: Designing for the Human-AI Team
Guido Gigante, Istituto Superiore di Sanità (ISS), Roma, Italy
Heatmaps and saliency maps have become the default tools for Explainable AI (XAI) in mammography, offering visually convincing highlights of suspicious regions. However, this reliance on visual plausibility creates a critical “Plausibility Trap.” While developers often equate these visually coherent outputs with true interpretability, human-centered studies reveal a starkly different clinical reality: providing a plausible heatmap does not inherently prevent diagnostic errors or mitigate automation bias when an AI system makes an incorrect suggestion.
This talk will explore the intersection of XAI and human factors through the lens of the Co-12 evaluation framework. Building on the foundations of technical correctness, we will shift the focus to the “User” dimensions of explanations—specifically Coherence, Context, and Controllability. We will discuss the illusion of visual overlap, examine what happens to the radiologist-in-the-loop when AI fails, and highlight why evaluating the algorithm in isolation is no longer sufficient. Ultimately, we will outline how to expand our evaluation landscape to actively design AI systems—such as prototype networks utilizing familiar clinical concepts—that foster a safer and more collaborative human-AI diagnostic team.
Explainability of AI Predictions: Perspectives from Breast Imaging Specialists
Aleksander Sadikovv, Faculty of Computer and Information Science, University of Ljubljana, Ljubljana, Slovenia
In this talk we will consider different notions of explainability, namely that it cannot be separated from the specific use case at hand, whom it is aimed at, and what the purpose of explanation is. We will discuss how breast imaging specialists perceive explainability and (un)certainty of AI predictions now and in the future. Finally, we will examine the relationship between concrete measures like Jaccard similarity coefficient on concrete examples and the subjective assessment of goodness of explanations on those examples by domain experts.
