2 papers
cs.CL2026
A Study on Hidden Layer Distillation for Large Language Model Pre-Training
Maxime Guigon, Lucas Dixon, Michaël E. Sander
Knowledge Distillation (KD) is a critical tool for training Large Language Models (LLMs), yet the majority of research focuses on approaches that rely solely on output logits, negl…
cs.CL2025
Improving Neutral Point-of-View Generation with Data- and Parameter-Efficient RL
Jessica Hoffmann, Christiane Ahlheim, Zac Yu +8
The paper shows that parameter-efficient reinforcement learning (PE-RL) is a highly effective training regime to improve large language models' (LLMs) ability to answer queries on…