1 paper · 1 filter
Semih Kara, OÄuzhan Ersoy
Conditioning a language model on additional context, such as feedback on a previous attempt, typically improves its response. Self-distillation trains the model to retain this impr…