papers
Publications (3)
cs.CL2022
Knowledge Distillation Transfer Sets and their Impact on Downstream NLU Tasks
Charith Peris, Lizhen Tan, Thomas Gueudre +3
Teacher-student knowledge distillation is a popular technique for compressing today's prevailing large language models into manageable sizes that fit low-latency downstream applica…
cs.CL2022
Alexa Teacher Model: Pretraining and Distilling Multi-Billion-Parameter Encoders for Natural Language Understanding Systems
Jack FitzGerald, Shankar Ananthakrishnan, Konstantine Arkoudas +38
We present results from a large-scale experiment on pretraining encoders with non-embedding parameter counts ranging from 700M to 9.3B, their subsequent distillation into smaller m…
cs.AI2026
Self-Improvement for Fast, High-Quality Plan Generation
Robert Gieselmann, Henrike von Huelsen, Mihai Samson +9
Generative models trained on synthetic plan data are a promising approach to generalized planning. Recent work has focused on finding any valid plan, rather than a high-quality sol…