1 citations · 1 across the 3 of their papers we have counts for
3 papers
FineWeb2: One Pipeline to Scale Them All -- Adapting Pre-Training Data Processing to Every Language
Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec +7
Pre-training state-of-the-art large language models (LLMs) requires vast amounts of clean and diverse text data. While the open development of large high-quality English pre-traini…
Analyzing & Reducing the Need for Learning Rate Warmup in GPT Training
Atli Kosson, Bettina Messmer, Martin Jaggi
Learning Rate Warmup is a popular heuristic for training neural networks, especially at larger batch sizes, despite limited understanding of its benefits. Warmup decreases the upda…
Towards an empirical understanding of MoE design choices
Dongyang Fan, Bettina Messmer, Martin Jaggi
In this study, we systematically evaluate the impact of common design choices in Mixture of Experts (MoEs) on validation performance, uncovering distinct influences at token and se…