22 citations · 42 across the 11 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2023★ 2 cited
Less or More From Teacher: Exploiting Trilateral Geometry For Knowledge Distillation
Chengming Hu, Haolun Wu, Xuan Li +5
Knowledge distillation aims to train a compact student network using soft supervision from a larger teacher network and hard supervision from ground truths. However, determining an…
cs.LG2022★ 10 cited
Gradient-based Bi-level Optimization for Deep Learning: A Survey
Can Chen, Xi Chen, Chen Ma +2
Bi-level optimization, especially the gradient-based category, has been widely used in the deep learning community including hyperparameter optimization and meta-knowledge extracti…