Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Predictable GRPO: A Closed-Form Model of Training Dynamics
Rajat Ghosh, Datta Nimmaturi, Aryan Singhal +4
We develop a first-principles reduced-order model of these dynamics. Under a single mean-field assumption that summarizes the policy by its expected reward, we reduce the GRPO upda…
cs.LG2025
BAR Conjecture: the Feasibility of Inference Budget-Constrained LLM Services with Authenticity and Reasoning
Jinan Zhou, Rajat Ghosh, Vaishnavi Bhargava +2
When designing LLM services, practitioners care about three key properties: inference-time budget, factual authenticity, and reasoning capacity. However, our analysis shows that no…