← Writing

How Do You Retrieve Medical Images by Lesion Region, Not Just the Whole Scan?

TL;DR — Give the retriever two experts. HMAR is a mixture-of-experts retrieval framework: Expert0 encodes the whole image for holistic similarity, while Expert1 learns position-invariant local features for lesion-region retrieval. A two-stage contrastive strategy removes the need for bounding-box annotations, a sliding-window matcher compares regions densely, and KAN layers produce compact hash codes. On RadioImageNet-CT it reaches 0.711 / 0.724 mAP at 64 / 128 bits.

Medical image retrieval (MIR) supports computer-aided diagnosis by finding past cases similar to the current one. Existing systems share three limitations:

Two experts, one retriever

HMAR (Hierarchical Modality-Aware Expert and Dynamic Routing) is built on a Mixture-of-Experts architecture:

No bounding boxes, fast search

Results

On RadioImageNet-CT (16 clinical patterns, 29,903 images), HMAR reaches 0.711 mAP with 64-bit codes and 0.724 mAP with 128-bit codes, improving over the state-of-the-art ACIR method by 0.7% and 1.1%.

Frequently asked questions

What is HMAR?

HMAR is a mixture-of-experts medical image retrieval framework with a global expert for whole-image similarity and a local expert for lesion-region retrieval, using KAN layers to generate hash codes.

Can medical image retrieval work without bounding-box labels?

Yes. HMAR uses a two-stage contrastive learning strategy that learns lesion-level representations without bounding-box annotations.

How well does HMAR perform?

On RadioImageNet-CT (16 clinical patterns, 29,903 images), HMAR achieves 0.711 mAP with 64-bit and 0.724 mAP with 128-bit hash codes, 0.7% and 1.1% above ACIR.


Based on HMAR: Hierarchical Modality-Aware Expert and Dynamic Routing Medical Image Retrieval Architecture (arXiv 2026, arXiv:2603.16679) by Aojie Yuan. Written by Aojie (Justin) Yuan, USC Fortis Lab.