2 papers
cs.LG2026
IsoCompute Playbook: Optimally Scaling Sampling Compute for LLM RL
Zhoujun Cheng, Yutao Xie, Yuxiao Qu +12
While scaling laws guide compute allocation for LLM pre-training, analogous prescriptions for reinforcement learning (RL) post-training of large language models (LLMs) remain poorl…
cs.AI2025
Flow of Reasoning: Training LLMs for Divergent Reasoning with Minimal Examples
Fangxu Yu, Lai Jiang, Haoqiang Kang +2
The ability to generate diverse solutions to a given problem is a hallmark of human creativity. This divergent reasoning is also crucial for machines, enhancing their robustness an…