1 paper · 1 filter
Jenny Schmalfuss, Nadine Chang, Vibashan VS +3
Vision language models (VLMs) respond to user-crafted text prompts and visual inputs, and are applied to numerous real-world problems. VLMs integrate visual modalities with large l…