1 paper
Chusen Li, Zhou Liu, Shuigeng Zhou +1
Large language models increasingly rely on either reinforcement learning or multi-agent prompting to improve reasoning, yet these two paradigms remain difficult to combine. Directl…