2 papers
cs.CL2022
Hidden State Variability of Pretrained Language Models Can Guide Computation Reduction for Transfer Learning
Shuo Xie, Jiahao Qiu, Ankita Pasad +3
While transferring a pretrained language model, common approaches conventionally attach their task-specific classifiers to the top layer and adapt all the pretrained layers. We inv…
cs.CL2019
The Curious Case of Neural Text Degeneration
Ari Holtzman, Jan Buys, Li Du +2
Despite considerable advancements with deep neural language models, the enigma of neural text degeneration persists when these models are tested as text generators. The counter-int…