1 paper · 1 filter
Xuan Zhang, Ruixiao Li, Zhijian Zhou +7
Reinforcement Learning (RL) has become a compelling way to strengthen the multi step reasoning ability of Large Language Models (LLMs). However, prevalent RL paradigms still lean o…