1 paper
Channe Chwa, Xinle Wu, Yao Lu
LLM post-training pipelines that combine supervised fine-tuning and reinforcement learning are difficult to configure under realistic compute budgets: the configuration space is hi…