ICCV 2025

Is Visual in-Context Learning for Compositional Medical Tasks within Reach?

S. Reiß, Z. Marinov, A. Jaus, C. Seibold, M. S. Sarfraz, E. Rodner, R. Stiefelhagen

IEEE/CVF International Conference on Computer Vision (ICCV), 2025

Medical Imaging Foundation Models

Key contributions

  • A synthetic compositional task generation engine that bootstraps sequential task chains from arbitrary segmentation datasets for training in-context learners.
  • A systematic analysis of codebook limitations in visual in-context learning and proposed improvements for capturing diverse task outputs.
  • Investigation of masking-based training objectives and their effect on compositional task generalization in medical imaging.
  • Important insights into multi-modal medical task sequences and the challenges that remain open for compositional in-context learning.
The compositional in-context learning pipeline: task sequences (super-resolution → inpainting → segmentation) are bootstrapped synthetically and processed by a single general model at test time, eliminating the need for multiple specialized models.
The compositional in-context learning pipeline: task sequences (super-resolution → inpainting → segmentation) are bootstrapped synthetically and processed by a single general model at test time, eliminating the need for multiple specialized models.

How it works

In this paper, we explore the potential of visual in-context learning to enable a single model to handle multiple tasks and adapt to new tasks during test time without re-training. Unlike previous approaches, our focus is on training in-context learners to adapt to sequences of tasks, rather than individual tasks. Our goal is to solve complex tasks that involve multiple intermediate steps using a single model, allowing users to define entire vision pipelines flexibly at test time.

To achieve this, we first examine the properties and limitations of visual in-context learning architectures, with a particular focus on the role of codebooks. We then introduce a novel method for training in-context learners using a synthetic compositional task generation engine. This engine bootstraps task sequences from arbitrary segmentation datasets, enabling the training of visual in-context learners for compositional tasks. Additionally, we investigate different masking-based training objectives to gather insights into how to train models better for solving complex, compositional tasks. Our exploration not only provides important insights especially for multi-modal medical task sequences but also highlights challenges that need to be addressed.

Citation

Reiß, S., Marinov, Z., Jaus, A., Seibold, C., Sarfraz, M. S., Rodner, E., & Stiefelhagen, R. (2025). Is Visual in-Context Learning for Compositional Medical Tasks within Reach? In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV 2025).