3 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
Zhuofeng Li, Haoxiang Zhang, Seungju Han +6
Outcome-driven reinforcement learning has advanced reasoning in large language models (LLMs), but prevailing tool-augmented approaches train a single, monolithic policy that interl…