88 citations · 114 across the 13 of their papers we have counts for
1 paper · 2 filters
Zhaohui Liang, Sivaramakrishnan Rajaraman, Niccolo Marini +2
CLIP and BiomedCLIP are examples of vision-language foundation models and offer strong cross-modal embeddings; however, they are not optimized for fine-grained medical retrieval ta…