Share:
Peer-Reviewed Publication
NPJ Digit Med2025;8(1):259.May 8, 2025Journal Article

Simulating mismatch between calibration and target population in AI for mammography the retrospective VAIB study.

Haiko Schurz1, Klara Solander2, Davida Åström2, Fernando Cossío2,3, Taeyang Choi2, Magnus Dustler4, Claes Lundström5,6, Håkan Gustafsson5,7, Sophia Zackrisson4,8, Fredrik Strand9,10
1Department of Oncology-Pathology, Karolinska Institutet, Solna, Sweden. haiko.schurz@ki.se.
2Department of Oncology-Pathology, Karolinska Institutet, Solna, Sweden.
3Medical Diagnostics Karolinska, Karolinska University Hospital, Solna, Sweden.
4Department of Translational Medicine, Diagnostic Radiology, Lund University, Malmö, Sweden.
5Center for Medical Image Science and Visualization (CMIV), Linköping University, Linköping, Sweden.
6Sectra AB, Linköping, Sweden.
7Department of Medical Radiation Physics, and Department of Health, Medicine and Caring Sciences, Linköping University, Linköping, Sweden.
8Department of Imaging and Physiology, Skåne University Hospital Malmö, Malmö, Sweden.
9Department of Oncology-Pathology, Karolinska Institutet, Solna, Sweden. fredrik.strand@ki.se.
10Medical Diagnostics Karolinska, Karolinska University Hospital, Solna, Sweden. fredrik.strand@ki.se.

Abstract

AI cancer detection models require calibration to attain the desired balance between cancer detection rate (CDR) and false positive rate. In this study, we simulate the impact of six types of mismatches between the calibration population and the clinical target population, by creating purposefully non-representative datasets to calibrate AI for clinical settings. Mismatching the acquisition year b…

Create a free account to keep reading

Free members get 10 full research views every month across publications, clinical trials, FDA clearances, adverse events, and NIH grants. No credit card required.

Want unlimited research access? See Pro plans

Data Accuracy Notice: Research intelligence on Health AI Central is aggregated from public sources (PubMed, ClinicalTrials.gov, FDA, NIH, CMS, and others) and refreshed nightly. Classifications and derived metrics are produced by automated methods described in our Methodology. We recommend verifying critical data points against the primary sources before making decisions.