2.7k citations · 2.7k across the 3 of their papers we have counts for
3 papers
cs.CL2023★ 2.7k cited
Llama 2: Open Foundation and Fine-Tuned Chat Models
Hugo Touvron, Louis Martin, Kevin Stone +65
In this work, we develop and release Llama 2, a collection of pretrained and fine-tuned large language models (LLMs) ranging in scale from 7 billion to 70 billion parameters. Our f…
cs.LG2023★ 2 cited
A Theory on Adam Instability in Large-Scale Machine Learning
Igor Molybog, Peter Albert, Moya Chen +14
We present a theory for the previously unexplained divergent behavior noticed in the training of large language models. We argue that the phenomenon is an artifact of the dominant…
cs.CL2021
Simple Local Attentions Remain Competitive for Long-Context Tasks
Wenhan Xiong, Barlas Oğuz, Anchit Gupta +5
Many NLP tasks require processing long contexts beyond the length limit of pretrained models. In order to scale these models to longer text sequences, many efficient long-range att…