1 paper
Stephen Ekaputra Limantoro, Jhe-Hao Lin, Chih-Yu Wang +4
Knowledge distillation (KD) compresses the network capacity by transferring knowledge from a large (teacher) network to a smaller one (student). It has been mainstream that the tea…