1 paper
Abbas Ghaddar, Ivan Kobyzev, Boxing Chen +1
Post-training hybridization of large language models (LLMs) often replaces quadratic self-attention with sliding-window attention (SWA) to reduce KV cache usage and improve latency…