3 papers
cs.LG2026
Aurora: A Leverage-Aware Spectral Optimizer
Alec Dewulf, Dhruv Pai, Li Yang +2
We show that for tall matrix parameters, like projection matrices in the MLP layers, the Muon update can have row norms that are arbitrarily non-uniform. This can lead to a self-re…
cs.AR2026
DataGuard: Guaranteeing Private Training in Systolic-array Based Accelerators
Pawan Kumar Sanjaya, Christina Giannoula, Nikhil Shreekumar +6
Differential privacy (DP) and federated learning (FL) have emerged as important privacy-preserving approaches when using sensitive data to train machine learning (ML) models. FL en…
cs.LG2026
Parallax: Parameterized Local Linear Attention for Language Modeling
Yifei Zuo, Dhruv Pai, Zhichen Zeng +3
Large Language Models (LLMs) have become the central paradigm in artificial intelligence, yet the core computational primitive of attention has remained structurally unchanged. Loc…