1 paper · 1 filter
Kengo Nakata, Daisuke Miyashita, Youyang Ng +2
In this paper, we rethink sparse lexical representations for image retrieval. By utilizing multi-modal large language models (M-LLMs) that support visual prompting, we can extract…