Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
On the Hidden Objective Biases of Group-based Reinforcement Learning
Aleksandar Fontana, Marco Simoni, Giulio Rossolini +2
Group-based reinforcement learning methods, like Group Relative Policy Optimization (GRPO), are widely used nowadays to post-train large language models. Despite their empirical su…
cs.LG2025
GTPO: Stabilizing Group Relative Policy Optimization via Gradient and Entropy Control
Marco Simoni, Aleksandar Fontana, Giulio Rossolini +2
Group Relative Policy Optimization (GRPO) is a promising policy-based approach for Large Language Model alignment, yet its performance is often limited by training instability and…