1 paper · 1 filter
Yirui Liu, Ruoling Qi, Xuaner Wu +2
Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. This structure complicates prefix caching:…