Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets
Medical Image Computing and Computer Assisted Intervention (MICCAI), 2026

Key contributions
A systematic benchmark spanning 4 datasets, 7 pseudo-label generators (nnU-Net, MedSAM, TotalSegmentator, STU-Net variants), and two evaluation scenarios — in-domain deployment vs. pre-train then fine-tune.
A demonstration that label quality is critical for in-domain deployment (strong positive correlation with model performance) but essentially irrelevant for pre-training: even the lowest-quality MedSAM pseudo-labels yield pre-trained models that match those trained on near-perfect labels after fine-tuning.
Actionable guidance for dataset creators: expert annotation effort is best invested in well-curated downstream target datasets, not in exhaustive refinement of massive pre-training corpora.
Manually refining radiological segmentation masks is highly resource-intensive. To determine when this expert commitment is truly justified for the training of segmentation models, we investigate the relationship between label quality and model performance. Expanding beyond models trained directly for inference, we conduct the first study isolating the impact of label quality in pre-training datasets.
While high-quality labels remain essential for models proceeding directly to deployment, we find no evidence that strict label quality is crucial for pre-training efficacy. These results question the necessity of exhaustive human-in-the-loop refinement for massive corpora intended for pre-training and suggest that expert effort is more effectively invested in well-curated downstream target datasets.
Citation
Jaus, A., Marinov, Z., Reiß, S., Seibold, C., Wei, J., Kleesiek, J., & Stiefelhagen, R. (2026). Good Enough? An Investigation on the Impact of Label Quality in Large-Scale Medical Datasets. Medical Image Computing and Computer Assisted Intervention (MICCAI).
