2 papers
cs.LG2026
Distill on a Diet: Efficient Knowledge Distillation via Learnable Data Pruning
Yifan Wu, Yiqi Wang, Xichen Ye +5
Knowledge Distillation (KD) is widely used to obtain compact models for efficient inference in resource-constrained environments. Yet the computational overhead of the distillation…
cs.LG2024
Optimized Gradient Clipping for Noisy Label Learning
Xichen Ye, Yifan Wu, Weizhong Zhang +3
Previous research has shown that constraining the gradient of loss function with respect to model-predicted probabilities can enhance the model robustness against noisy labels. The…