3 papers
cs.CV2026
Watch Before You Answer: Learning from Visually Grounded Post-Training
Yuxuan Zhang, EunJeong Hwang, Huaisong Zhang +8
It is critical for vision-language models (VLMs) to comprehensively understand visual, temporal, and textual cues. However, despite rapid progress in multimodal modeling, video und…
cs.CV2025
The in-context inductive biases of vision-language models differ across modalities
Kelsey Allen, Ishita Dasgupta, Eliza Kosoy +1
Inductive biases are what allow learners to make guesses in the absence of conclusive evidence. These biases have often been studied in cognitive science using concepts or categori…
cs.CV2025
Decoupling the components of geometric understanding in Vision Language Models
Eliza Kosoy, Annya Dahmani, Andrew K. Lampinen +4
Understanding geometry relies heavily on vision. In this work, we evaluate whether state-of-the-art vision language models (VLMs) can understand simple geometric concepts. We use a…