1 paper · 1 filter
Konstantin Berestizshevsky, Renzo Andri, Lukas Cavigelli
We present Top-Theta (Top-θ) Attention, a training-free method for sparsifying transformer attention during inference. Our key insight is that static, per-head thresholds can be…