1 citations · 1 across the 2 of their papers we have counts for
3 papers
On Difficulties of Attention Factorization through Shared Memory
Uladzislau Yorsh, Martin Holeňa, Ondřej Bojar +1
Transformers have revolutionized deep learning in numerous fields, including natural language processing, computer vision, and audio processing. Their strength lies in their attent…
Linear Self-Attention Approximation via Trainable Feedforward Kernel
Uladzislau Yorsh, Alexander Kovalenko
In pursuit of faster computation, Efficient Transformers demonstrate an impressive variety of approaches -- models attaining sub-quadratic attention complexity can utilize a notion…
SimpleTRON: Simple Transformer with O(N) Complexity
Uladzislau Yorsh, Alexander Kovalenko, Vojtěch Vančura +3
In this paper, we propose that the dot product pairwise matching attention layer, which is widely used in Transformer-based models, is redundant for the model performance. Attentio…