1 paper
Zhongzhu Zhou, Fengxiang Bie, Ziyan Chen +6
Converting pretrained attention modules such as grouped-query attention (GQA) into multi-head latent attention (MLA) can improve expressivity without increasing KV-cache cost, maki…