Publications (4)
A Study on Knowledge Distillation from Weak Teacher for Scaling Up Pre-trained Language Models
Hayeon Lee, Rui Hou, Jongpil Kim +3
Distillation from Weak Teacher (DWT) is a method of transferring knowledge from a smaller, weaker teacher model to a larger student model to improve its performance. Previous studi…
FRESCO: Benchmarking and Optimizing Re-rankers for Evolving Semantic Conflict in Retrieval-Augmented Generation
Sohyun An, Hayeon Lee, Shuibenyang Yuan +4
Retrieval-Augmented Generation (RAG) is a key approach to mitigating the temporal staleness of large language models (LLMs) by grounding responses in up-to-date evidence. Within th…
Cycle-Consistent Search: Question Reconstructability as a Proxy Reward for Search Agent Training
Sohyun An, Shuibenyang Yuan, Hayeon Lee +2
Reinforcement Learning (RL) has shown strong potential for optimizing search agents in complex information retrieval tasks. However, existing approaches predominantly rely on gold…
Co-training and Co-distillation for Quality Improvement and Compression of Language Models
Hayeon Lee, Rui Hou, Jongpil Kim +4
Knowledge Distillation (KD) compresses computationally expensive pre-trained language models (PLMs) by transferring their knowledge to smaller models, allowing their use in resourc…