5 citations · 5 across the 9 of their papers we have counts for
1 paper · 2 filters
Megh Thakkar, Quentin Fournier, Matthew D Riemer +4
Large language models are first pre-trained on trillions of tokens and then instruction-tuned or aligned to specific preferences. While pre-training remains out of reach for most r…