3 papers
cs.LG2025
Dion2: A Simple Method to Shrink Matrix in Muon
Kwangjun Ahn, Noah Amsel, John Langford
The Muon optimizer enjoys strong empirical performance and theoretical grounding. However, the super-linear cost of its orthonormalization step introduces increasing overhead with…
cs.LG2025
Dion: Distributed Orthonormalized Updates
Kwangjun Ahn, Byron Xu, Natalie Abreu +5
Orthonormalized updates accelerate training, improve stability, and enable robust hyperparameter transfer, but existing methods like Muon rely on dense matrix operations that clash…
cs.LG2025
Efficient Joint Prediction of Multiple Future Tokens
Kwangjun Ahn, Alex Lamb, John Langford
In this short report, we introduce joint multi-token prediction (JTP), a lightweight modification of standard next-token prediction designed to enrich hidden state representations…