195 citations · 261 across the 8 of their papers we have counts for
Showing 2020Show all
3 papers · 1 filter
cs.LG2020★ 195 cited
Long Range Arena: A Benchmark for Efficient Transformers
Yi Tay, Mostafa Dehghani, Samira Abnar +7
Transformers do not scale very well to long sequence lengths largely because of quadratic self-attention complexity. In the recent months, a wide spectrum of efficient, fast Transf…
cs.LG2020★ 13 cited
Transferring Inductive Biases through Knowledge Distillation
Samira Abnar, Mostafa Dehghani, Willem Zuidema
Having the right inductive biases can be crucial in many tasks or scenarios where data or computing resources are a limiting factor, or where training data is not perfectly represe…
cs.LG2020
Quantifying Attention Flow in Transformers
Samira Abnar, Willem Zuidema
In the Transformer model, "self-attention" combines information from attended embeddings into the representation of the focal embedding in the next layer. Thus, across layers of th…