2 papers
cs.CL2025
Essential-Web v1.0: 24T tokens of organized web data
Essential AI, :, Andrew Hojel +22
Data plays the most prominent role in how language models acquire skills and knowledge. The lack of massive, well-organized pre-training datasets results in costly and inaccessible…
cs.LG2025
Practical Efficiency of Muon for Pretraining
Essential AI, :, Ishaan Shah +22
We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon…