1 paper · 1 filter
Zhaohui Liang, Sivaramakrishnan Rajaraman, Niccolo Marini +2
CLIP and BiomedCLIP are examples of vision-language foundation models and offer strong cross-modal embeddings; however, they are not optimized for fine-grained medical retrieval ta…