1 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.LG2023
LABO: Towards Learning Optimal Label Regularization via Bi-level Optimization
Peng Lu, Ahmad Rashid, Ivan Kobyzev +2
Regularization techniques are crucial to improving the generalization performance and training efficiency of deep neural networks. Many deep learning algorithms rely on weight deca…
cs.CL2022★ 1 cited
Improving Generalization of Pre-trained Language Models via Stochastic Weight Averaging
Peng Lu, Ivan Kobyzev, Mehdi Rezagholizadeh +3
Knowledge Distillation (KD) is a commonly used technique for improving the generalization of compact Pre-trained Language Models (PLMs) on downstream tasks. However, such methods i…
cs.LG2022★ 1 cited
Do we need Label Regularization to Fine-tune Pre-trained Language Models?
Ivan Kobyzev, Aref Jafari, Mehdi Rezagholizadeh +5
Knowledge Distillation (KD) is a prominent neural model compression technique that heavily relies on teacher network predictions to guide the training of a student model. Consideri…