1 paper · 1 filter
Pooyan Rahmanzadehgervi, Logan Bolton, Mohammad Reza Taesiri +1
While large language models with vision capabilities (VLMs), e.g., GPT-4o and Gemini 1.5 Pro, score high on many vision-understanding benchmarks, they are still struggling with low…