3 papers
cs.LG2026
MoE Lens -- An Expert Is All You Need
Marmik Chaudhari, Idhant Gulati, Nishkal Hundia +2
Mixture of Experts (MoE) models enable parameter-efficient scaling through sparse expert activations, yet optimizing their inference and memory costs remains challenging due to lim…
cs.LG2026
Sparse Crosscoders for diffing MoEs and Dense models
Marmik Chaudhari, Nishkal Hundia, Idhant Gulati
Mixture of Experts (MoE) achieve parameter-efficient scaling through sparse expert routing, yet their internal representations remain poorly understood compared to dense models. We…
cs.LG2025
Sparsity and Superposition in Mixture of Experts
Marmik Chaudhari, Jeremi Nuer, Rome Thorstenson
Mixture of Experts (MoE) models have become central to scaling large language models, yet their mechanistic differences from dense networks remain poorly understood. Previous work…