1 paper · 1 filter
Jiabin Fan, Guoqing Luo, Michael Bowling +1
We propose a novel k-step return estimation method (called KETCHUP) for Reinforcement Learning(RL)-based knowledge distillation (KD) in text generation tasks. Our idea is to induce…