1 paper
Dimitri von Rütte, Janis Fluri, Omead Pooladzandi +3
Modern LLM pre-training consumes vast amounts of compute and training data, making the scaling behavior, or scaling laws, of different models a key distinguishing factor. Discrete…