1 paper
Disen Lan, Jianbin Zheng, Yuxi Ren +5
Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the e…