MICCAI 2026

Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets

A. Jaus, Z. Marinov, S. Reiß, C. Seibold, J. Wei, J. Kleesiek, R. Stiefelhagen

Medical Image Computing and Computer Assisted Intervention (MICCAI), 2026

Medical Imaging Datasets
The same abdominal CT scan segmented by four different automated pseudo-labelers (one per panel) — each a plausible but visibly different labeling of the liver (red), kidneys (yellow and brown), pancreas (dark green), and spleen (light green). This annotator disagreement is exactly the label-quality variation the paper studies.
The same abdominal CT scan segmented by four different automated pseudo-labelers (one per panel) — each a plausible but visibly different labeling of the liver (red), kidneys (yellow and brown), pancreas (dark green), and spleen (light green). This annotator disagreement is exactly the label-quality variation the paper studies.

Key contributions

01

A systematic benchmark spanning 4 datasets, 7 pseudo-label generators (nnU-Net, MedSAM, TotalSegmentator, STU-Net variants), and two evaluation scenarios — in-domain deployment vs. pre-train then fine-tune.

02

A demonstration that label quality is critical for in-domain deployment (strong positive correlation with model performance) but essentially irrelevant for pre-training: even the lowest-quality MedSAM pseudo-labels yield pre-trained models that match those trained on near-perfect labels after fine-tuning.

03

Actionable guidance for dataset creators: expert annotation effort is best invested in well-curated downstream target datasets, not in exhaustive refinement of massive pre-training corpora.

Manually refining radiological segmentation masks is highly resource-intensive. To determine when this expert commitment is truly justified for the training of segmentation models, we investigate the relationship between label quality and model performance. Expanding beyond models trained directly for inference, we conduct the first study isolating the impact of label quality in pre-training datasets.

While high-quality labels remain essential for models proceeding directly to deployment, we find no evidence that strict label quality is crucial for pre-training efficacy. These results question the necessity of exhaustive human-in-the-loop refinement for massive corpora intended for pre-training and suggest that expert effort is more effectively invested in well-curated downstream target datasets.

Citation

Jaus, A., Marinov, Z., Reiß, S., Seibold, C., Wei, J., Kleesiek, J., & Stiefelhagen, R. (2026). Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets. Medical Image Computing and Computer Assisted Intervention (MICCAI).