9 citations · 22 across the 8 of their papers we have counts for
Showing cs.CVShow all
2 papers · 1 filter
cs.CV2024★ 1 cited
CROME: Cross-Modal Adapters for Efficient Multimodal LLM
Sayna Ebrahimi, Sercan O. Arik, Tejas Nama +1
Multimodal Large Language Models (MLLMs) demonstrate remarkable image-language capabilities, but their widespread use faces challenges in cost-effective training and adaptation. Ex…
cs.CV2023★ 2 cited
Pic2Word: Mapping Pictures to Words for Zero-shot Composed Image Retrieval
Kuniaki Saito, Kihyuk Sohn, Xiang Zhang +4
In Composed Image Retrieval (CIR), a user combines a query image with text to describe their intended target. Existing methods rely on supervised learning of CIR models using label…