1 paper
Reilly Haskins, Benjamin Adams
Knowledge distillation compresses a larger neural model (teacher) into smaller, faster student models by training them to match teacher outputs. However, the internal computational…