1 citations · 1 across the 15 of their papers we have counts for
1 paper · 1 filter
Sushant Mehta, Raj Dandekar, Rajat Dandekar +1
We present the first comprehensive study of latent multi-head attention (MLA) for small language models, revealing interesting efficiency-quality trade-offs. Training 30M-parameter…