1 paper
Boren Hu, Yun Zhu, Jiacheng Li +1
Dynamic early exiting has been proven to improve the inference speed of the pre-trained language model like BERT. However, all samples must go through all consecutive layers before…