1 paper · 1 filter
Daniel Goldstein, Eric Alcaide, Janna Lu +1
We present Rapid Attention Distillation to Linear Attention Decoders at Scale (RADLADS), a protocol for rapidly converting softmax attention transformers into linear attention deco…