2 citations · 2 across the 6 of their papers we have counts for
1 paper · 1 filter
Mingze Wu, Abhinav Anand, Shweta Verma +1
Post-training using online reinforcement learning (RL) is an important training step for LLMs, including code-generating models. However, online RL for code generation involves LLM…