1 paper
Wenxuan Jiang, Zining Fan, Zijian Zhang +6
Reinforcement Learning (RL) has enabled LLMs to excel in objective reasoning tasks such as mathematics and code generation. However, applying RL to open-ended tasks, such as creati…