12 citations · 12 across the 1 of their papers we have counts for
1 paper
Kshitij Gupta, Benjamin Thérien, Adam Ibrahim +5
Large language models (LLMs) are routinely pre-trained on billions of tokens, only to restart the process over again once new data becomes available. A much cheaper and more effici…