1.8k citations · 1.8k across the 3 of their papers we have counts for
1 paper · 1 filter
Yanke Zhou, Yiduo Li, Hanlin Tang +6
Long-context inference in large language models is bottlenecked by the quadratic cost of full attention. Existing efficient alternatives often rely either on native sparse training…