13 citations · 15 across the 5 of their papers we have counts for
1 paper · 1 filter
Tianle Zhong, Neiwen Ling, Yifan Pi +5
Modern LLM RL systems separate rollout generation from policy optimization. These two stages are expected to produce token probabilities that match exactly. However, implementation…