4 papers
Representations of Text and Images Align From Layer One
Evžen Wybitul, Javier Rando, Florian Tramèr +1
We show that for a variety of concepts in adapter-based vision-language models, the representations of their images and their text descriptions are meaningfully aligned from the ve…
Access Controls Will Solve the Dual-Use Dilemma
Evžen Wybitul
AI safety systems face the dual-use dilemma. It is unclear whether to answer dual-use requests, since the same query could be either harmless or harmful depending on who made it an…
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
Alex Cloud, Jacob Goldman-Wetzler, Evžen Wybitul +2
Neural networks are trained primarily based on their inputs and outputs, without regard for their internal mechanisms. These neglected mechanisms determine properties that are crit…
ViSTa Dataset: Do vision-language models understand sequential tasks?
Evžen Wybitul, Evan Ryan Gunter, Mikhail Seleznyov +1
Using vision-language models (VLMs) as reward models in reinforcement learning holds promise for reducing costs and improving safety. So far, VLM reward models have only been used…