1 paper
Meng Cao, Lei Shu, Lei Yu +4
Reinforcement learning (RL) can align language models with non-differentiable reward signals, such as human preferences. However, a major challenge arises from the sparsity of thes…