164 citations · 413 across the 20 of their papers we have counts for
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2023★ 14 cited
D4: Improving LLM Pretraining via Document De-Duplication and Diversification
Kushal Tirumala, Daniel Simig, Armen Aghajanyan +1
Over recent years, an increasing amount of compute and data has been poured into training large language models (LLMs), usually by doing one-pass learning on as many tokens as poss…
cs.CL2020
Reservoir Transformers
Sheng Shen, Alexei Baevski, Ari S. Morcos +3
We demonstrate that transformers obtain impressive performance even when some of the layers are randomly initialized and never updated. Inspired by old and well-established ideas i…