Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Representations of Text and Images Align From Layer One
Evžen Wybitul, Javier Rando, Florian Tramèr +1
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…
cs.CV2024
ViSTa Dataset: Do vision-language models understand sequential tasks?
Evžen Wybitul, Evan Ryan Gunter, Mikhail Seleznyov +1
Using vision-language models (VLMs) as reward models in reinforcement learning holds promise for reducing costs and improving safety. So far, VLM reward models have only been used…