1 paper
Changhai Zhou, Kieran Liu, Yuhua Zhou +17
Long-context RL post-training is constrained by the lifetime of state and gradients, not attention cost alone. In GRPO, one multi-million-token prompt must serve old-policy and ref…