3 papers
cs.CV2024
ViSTa Dataset: Do vision-language models understand sequential tasks?
Evžen Wybitul, Evan Ryan Gunter, Mikhail Seleznyov +1
Using vision-language models (VLMs) as reward models in reinforcement learning holds promise for reducing costs and improving safety. So far, VLM reward models have only been used…
cs.LG2024
NGD converges to less degenerate solutions than SGD
Moosa Saghir, N. R. Raghavendra, Zihe Liu +1
The number of free parameters, or dimension, of a model is a straightforward way to measure its complexity: a model with more parameters can encode more information. However, this…
cs.AI2024
Quantifying stability of non-power-seeking in artificial agents
Evan Ryan Gunter, Yevgeny Liokumovich, Victoria Krakovna
We investigate the question: if an AI agent is known to be safe in one setting, is it also safe in a new setting similar to the first? This is a core question of AI alignment--we t…