1 paper
Giovanni De Muri, Mark Vero, Robin Staab +1
LLMs are often used by downstream users as teacher models for knowledge distillation, compressing their capabilities into memory-efficient models. However, as these teacher models…