1 paper · 1 filter
Disen Lan, Jianbin Zheng, Yuxi Ren +5
Hybrid attention models improve long-context efficiency by retaining only a subset of full-attention layers and replacing the remaining layers with linear attention. However, the e…