autoPET3

The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT – Multitracer Multicenter Generalization

J. Dexl, K. Jeblick, A. Mittermeier, B. Schachtner, A. T. Stüber, J. Topalis, M. Rokuss, F. Isensee, K. H. Maier-Hein, H. Kalisch, J. Kleesiek, C. M. Seibold, H. Alasmawi, L. Y. L. Chan, Y. Yuan, A. Jaus, R. Stiefelhagen, P. O. Megne Choudjа, K. Nikolaou, C. La Fougère, S. Gatidis, M. P. Fabritius, M. Heimer, G. Abaci, L. K. Shiyam Sundar, R. A. Werner, J. Ricke, C. C. Cyran, T. Küstner, M. Ingrisch

MICCAI 2024 Challenge (autoPET3), 2026

Medical Imaging Segmentation Datasets

Key contributions

  • The largest publicly available annotated PSMA PET/CT dataset (597 studies from LMU Munich), enabling multitracer research.
  • A compositional generalization benchmark with four held-out tracer–center combinations, two of which were entirely unseen during training.
  • A complementary data-centric award category isolating data handling strategies from architecture choices with a fixed baseline model.
  • Systematic patient- and lesion-level analysis showing that case difficulty and heterogeneity dominate algorithm differences among top-ranked teams.
Representative test cases from the autoPET3 challenge illustrating the core difficulty: wide variation in disease extent (single lesion to high metastatic burden) and tracer-specific uptake patterns across PSMA (left) and FDG (right) scans from two centers.
Representative test cases from the autoPET3 challenge illustrating the core difficulty: wide variation in disease extent (single lesion to high metastatic burden) and tracer-specific uptake patterns across PSMA (left) and FDG (right) scans from two centers.

How it works

We report the design and results of the third autoPET challenge (MICCAI 2024), which benchmarked automated lesion segmentation in whole-body PET/CT under a compositional generalization setting. Training data comprised 1,014 [18F]-FDG PET/CT studies from the University Hospital Tübingen and 597 [18F]/[68Ga]-PSMA PET/CT studies from the LMU University Hospital Munich, constituting the largest publicly available annotated PSMA PET/CT dataset to date. The held-out test set of 200 studies covered four tracer–center combinations, two of which represented unseen compositional pairings.

Seventeen teams submitted 27 algorithms, predominantly nnU-Net-based 3D networks with PET/CT channel concatenation. The top-ranked algorithm achieved a mean DSC of 0.66, FNV of 3.18 mL, and FPV of 2.78 mL across all four test conditions, improving DSC by 8% and reducing the false-negative volume by 5 mL relative to the provided baseline. Three main conclusions can be drawn: (1) in-domain multitracer PET/CT segmentation is sufficient and probably approaching reader agreement; (2) compositional generalization to unseen tracer–center combinations remains an open problem mainly driven by systematic volume overestimation; (3) heterogeneity and case difficulty drive performance variation substantially more than the choice of algorithm among top-ranked teams.

Citation

Dexl, J., Jeblick, K., Mittermeier, A., Schachtner, B., Stüber, A. T., Topalis, J., Rokuss, M., Isensee, F., Maier-Hein, K. H., Kalisch, H., Kleesiek, J., Seibold, C. M., Alasmawi, H., Chan, L. Y. L., Yuan, Y., Jaus, A., Stiefelhagen, R., et al. (2026). The autoPET3 Challenge: Automated Lesion Segmentation in Whole-Body PET/CT – Multitracer Multicenter Generalization. MICCAI 2024 Challenge (autoPET3). arXiv:2605.05775.