8 citations · 8 across the 4 of their papers we have counts for
1 paper · 1 filter
Tao Ji, Bin Guo, Yuanbin Wu +6
Multi-head Latent Attention (MLA) is an innovative architecture proposed by DeepSeek, designed to ensure efficient and economical inference by significantly compressing the Key-Val…