2 papers
cs.LG2025
Practical Efficiency of Muon for Pretraining
Essential AI, :, Ishaan Shah +22
We demonstrate that Muon, the simplest instantiation of a second-order optimizer, explicitly expands the Pareto frontier over AdamW on the compute-time tradeoff. We find that Muon…
cs.CL2025
Rethinking Reflection in Pre-Training
Essential AI, :, Darsh J Shah +26
A language model's ability to reflect on its own reasoning provides a key advantage for solving complex problems. While most recent research has focused on how this ability develop…