1 paper · 1 filter
Plawan Kumar Rath
We show that knowledge distillation (KD) in small instruction-tuned language models has asymmetric effects on bias, and that measuring them correctly requires accounting for where…