2 papers
cs.LG2026
Rethinking Reverse KL as Adaptive Entropy Distillation
Shizhen Li, Zhiyu Shen, Yuyin Lu +4
Knowledge distillation (KD) is widely used to transfer the capabilities of large language models (LLMs) to smaller students, but existing objectives often struggle to balance faith…
cs.CL2024
IterSelectTune: An Iterative Training Framework for Efficient Instruction-Tuning Data Selection
Jielin Song, Siyu Liu, Bin Zhu +1
As large language models (LLMs) continue to advance, instruction tuning has become critical for improving their ability to generate accurate and contextually appropriate responses.…