1 paper · 1 filter
Weiye Shi, Fanxu Meng, Muhan Zhang
Multi-head latent attention (MLA) is increasingly important for long-context LLM inference because compact latent states replace the growing key-value (KV) cache and reduce decodin…