2 papers
cs.CL2025
KETCHUP: K-Step Return Estimation for Sequential Knowledge Distillation
Jiabin Fan, Guoqing Luo, Michael Bowling +1
We propose a novel k-step return estimation method (called KETCHUP) for Reinforcement Learning(RL)-based knowledge distillation (KD) in text generation tasks. Our idea is to induce…
cs.IR2024
Best Practices for Distilling Large Language Models into BERT for Web Search Ranking
Dezhi Ye, Junwei Hu, Jiabin Fan +4
Recent studies have highlighted the significant potential of Large Language Models (LLMs) as zero-shot relevance rankers. These methods predominantly utilize prompt learning to ass…