Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
GRPO is Secretly a Process Reward Model
Michael Sullivan, Alexander Koller
Process reward models (PRMs) allow for fine-grained credit assignment in reinforcement learning (RL), and seemingly contrast with outcome reward models (ORMs), which assign a singl…
cs.LG2025
Procedural Environment Generation for Tool-Use Agents
Michael Sullivan, Mareike Hartmann, Alexander Koller
Although the power of LLM tool-use agents has ignited a flurry of recent research in this area, the curation of tool-use training data remains an open problemespecially for onli…