7 citations · 22 across the 29 of their papers we have counts for
1 paper · 2 filters
Gregor Geigle, Abhay Jain, Radu Timofte +1
Modular vision-language models (Vision-LLMs) align pretrained image encoders with (frozen) large language models (LLMs) and post-hoc condition LLMs to `understand' the image input.…