1 paper
Feng Zhang, Zezhong Tan, Xinhong Ma +6
To address the limited capability expansion and low sample efficiency of Reinforcement Learning (RL), recent methods have integrated ''hints'' into post-training, which are prefix…