How Do You Retrieve Medical Images by Lesion Region, Not Just the Whole Scan?
Medical image retrieval (MIR) supports computer-aided diagnosis by finding past cases similar to the current one. Existing systems share three limitations:
- Uniform encoding that ignores the varying clinical importance of anatomical structures;
- Coarse similarity based on classification labels;
- Global-only matching, when clinicians often need cases with a similar lesion, not a similar-looking scan.
Two experts, one retriever
HMAR (Hierarchical Modality-Aware Expert and Dynamic Routing) is built on a Mixture-of-Experts architecture:
- Expert0 extracts global features for holistic similarity matching.
- Expert1 learns position-invariant local representations for precise lesion-region retrieval.
No bounding boxes, fast search
- A two-stage contrastive learning strategy removes the need for expensive bounding-box annotations.
- A sliding-window matching algorithm enables dense local comparison at inference time.
- Kolmogorov-Arnold Network (KAN) layers generate hash codes for efficient Hamming-distance search.
Results
On RadioImageNet-CT (16 clinical patterns, 29,903 images), HMAR reaches 0.711 mAP with 64-bit codes and 0.724 mAP with 128-bit codes, improving over the state-of-the-art ACIR method by 0.7% and 1.1%.
Frequently asked questions
What is HMAR?
HMAR is a mixture-of-experts medical image retrieval framework with a global expert for whole-image similarity and a local expert for lesion-region retrieval, using KAN layers to generate hash codes.
Can medical image retrieval work without bounding-box labels?
Yes. HMAR uses a two-stage contrastive learning strategy that learns lesion-level representations without bounding-box annotations.
How well does HMAR perform?
On RadioImageNet-CT (16 clinical patterns, 29,903 images), HMAR achieves 0.711 mAP with 64-bit and 0.724 mAP with 128-bit hash codes, 0.7% and 1.1% above ACIR.
Based on HMAR: Hierarchical Modality-Aware Expert and Dynamic Routing Medical Image Retrieval Architecture (arXiv 2026, arXiv:2603.16679) by Aojie Yuan. Written by Aojie (Justin) Yuan, USC Fortis Lab.