Learning Anatomical and Pathological Priors from Multi-modal and Report-based Data
Despite recent advances in biomedical image segmentation, most approaches lack explicit incorporation of anatomical knowledge. While radiologists rely on a structured understanding of the human body, models typically learn implicit representations focused on localized regions.
To address this gap, our project builds a dataset covering the full human body across multiple imaging modalities, and uses radiological reports as weak supervision to capture how observed anatomy aligns with — or deviates from — normative patterns.
Combining these sources, we train models that explicitly encode anatomical and pathological priors. One key application is semantic search across anatomical datasets, supporting queries such as:
“Retrieve all CT scans with left lungs of diameter X showing signs of pulmonary embolism.”

