2 papers
cs.LG2025
GUIDE: Guided Initialization and Distillation of Embeddings
Khoa Trinh, Gaurav Menghani, Erik Vee
Algorithmic efficiency techniques such as distillation (\cite{hinton2015distillation}) are useful in improving model quality without increasing serving costs, provided a larger tea…
cs.LG2024
Linear Projections of Teacher Embeddings for Few-Class Distillation
Noel Loo, Fotis Iliopoulos, Wei Hu +1
Knowledge Distillation (KD) has emerged as a promising approach for transferring knowledge from a larger, more complex teacher model to a smaller student model. Traditionally, KD i…