4 citations · 5 across the 2 of their papers we have counts for
2 papers
cs.CL2021★ 4 cited
A Short Study on Compressing Decoder-Based Language Models
Tianda Li, Yassir El Mesbahi, Ivan Kobyzev +6
Pre-trained Language Models (PLMs) have been successful for a wide range of natural language processing (NLP) tasks. The state-of-the-art of PLMs, however, are extremely large to b…
cs.CL2021★ 1 cited
RAIL-KD: RAndom Intermediate Layer Mapping for Knowledge Distillation
Md Akmal Haidar, Nithin Anchuri, Mehdi Rezagholizadeh +3
Intermediate layer knowledge distillation (KD) can improve the standard KD technique (which only targets the output of teacher and student models) especially over large pre-trained…