1 paper · 2 filters
Baohao Liao, David Thulke, Sanjika Hewavitharana +2
The pre-training of masked language models (MLMs) consumes massive computation to achieve good results on downstream NLP tasks, resulting in a large carbon footprint. In the vanill…