1 paper
Megh Thakkar, Quentin Fournier, Matthew D Riemer +4
Large language models are first pre-trained on trillions of tokens and then instruction-tuned or aligned to specific preferences. While pre-training remains out of reach for most r…