1 paper · 1 filter
Gregor Geigle, Abhay Jain, Radu Timofte +1
Modular vision-language models (Vision-LLMs) align pretrained image encoders with (frozen) large language models (LLMs) and post-hoc condition LLMs to `understand' the image input.…