2 citations · 3 across the 8 of their papers we have counts for
1 paper · 2 filters
Wenhan Ma, Jianyu Wei, Liang Zhao +10
Modern large language models (LLMs) rely on reinforcement learning during post-training to push specific capabilities, yet integrating multiple capabilities into one model remains…