1 paper · 1 filter
Reilly Haskins, Benjamin Adams
Knowledge distillation compresses a larger neural model (teacher) into smaller, faster student models by training them to match teacher outputs. However, the internal computational…