1 paper · 1 filter
Baohao Liao, David Thulke, Sanjika Hewavitharana +2
The pre-training of masked language models (MLMs) consumes massive computation to achieve good results on downstream NLP tasks, resulting in a large carbon footprint. In the vanill…