1 citations · 1 across the 3 of their papers we have counts for
3 papers
Muon: Training and Trade-offs with Latent Attention and MoE
Sushant Mehta, Raj Dandekar, Rajat Dandekar +1
We present a comprehensive theoretical and empirical study of the Muon optimizer for training transformers only with a small to medium decoder (30M - 200M parameters), with an emph…
Unifying Mixture of Experts and Multi-Head Latent Attention for Efficient Language Models
Sushant Mehta, Raj Dandekar, Rajat Dandekar +1
We present MoE-MLA-RoPE, a novel architecture combination that combines Mixture of Experts (MoE) with Multi-head Latent Attention (MLA) and Rotary Position Embeddings (RoPE) for ef…
Latent Multi-Head Attention for Small Language Models
Sushant Mehta, Raj Dandekar, Rajat Dandekar +1
We present the first comprehensive study of latent multi-head attention (MLA) for small language models, revealing interesting efficiency-quality trade-offs. Training 30M-parameter…