2 papers
cs.CV2024
ViSTa Dataset: Do vision-language models understand sequential tasks?
Evžen Wybitul, Evan Ryan Gunter, Mikhail Seleznyov +1
Using vision-language models (VLMs) as reward models in reinforcement learning holds promise for reducing costs and improving safety. So far, VLM reward models have only been used…
cs.LG2024
NGD converges to less degenerate solutions than SGD
Moosa Saghir, N. R. Raghavendra, Zihe Liu +1
The number of free parameters, or dimension, of a model is a straightforward way to measure its complexity: a model with more parameters can encode more information. However, this…