2.1k citations · 3.5k across the 14 of their papers we have counts for
Showing 2021Show all
2 papers · 1 filter
cs.LG2021★ 16 cited
Show Your Work: Scratchpads for Intermediate Computation with Language Models
Maxwell Nye, Anders Johan Andreassen, Guy Gur-Ari +9
Large pre-trained language models perform remarkably well on tasks that can be done "in one pass", such as generating realistic text or synthesizing computer programs. However, the…
cs.LG2021★ 16 cited
How to decay your learning rate
Aitor Lewkowycz
Complex learning rate schedules have become an integral part of deep learning. We find empirically that common fine-tuned schedules decay the learning rate after the weight norm bo…