1 paper · 1 filter
Guihong Li, Mehdi Rezagholizadeh, Mingyu Yang +2
Multi-head latent attention (MLA) is designed to optimize KV cache memory through low-rank key-value joint compression. Rather than caching keys and values separately, MLA stores t…