Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Constrained Group Relative Policy Optimization
Roger Girgis, Rodrigue de Schaetzen, Luke Rowe +3
While Group Relative Policy Optimization (GRPO) has emerged as a scalable framework for critic-free policy learning, extending it to settings with explicit behavioral constraints r…
cs.LG2025
Neural Coherence : Find higher performance to out-of-distribution tasks from few samples
Simon Guiroy, Mats Richter, Sarath Chandar +1
To create state-of-the-art models for many downstream tasks, it has become common practice to fine-tune a pre-trained large vision model. However, it remains an open question of ho…