1 paper
Hyemin Lim, Jaeyeon Lee, Dong-Wan Choi
Large pretrained language models such as BERT suffer from slow inference and high memory usage, due to their huge size. Recent approaches to compressing BERT rely on iterative prun…