1 paper · 1 filter
Leander Girrbach, Stephan Alaniz, Yiran Huang +2
Pre-trained large language models (LLMs) have been reliably integrated with visual input for multimodal tasks. The widespread adoption of instruction-tuned image-to-text vision-lan…