54 citations · 55 across the 3 of their papers we have counts for
3 papers
cs.CV2023
FFT-based Dynamic Token Mixer for Vision
Yuki Tatsunami, Masato Taki
Multi-head-self-attention (MHSA)-equipped models have achieved notable performance in computer vision. Their computational complexity is proportional to quadratic numbers of pixels…
cs.CV2022★ 54 cited
Sequencer: Deep LSTM for Image Classification
Yuki Tatsunami, Masato Taki
In recent computer vision research, the advent of the Vision Transformer (ViT) has rapidly revolutionized various architectural design efforts: ViT achieved state-of-the-art image…
cs.CV2021★ 1 cited
RaftMLP: How Much Can Be Done Without Attention and with Less Spatial Locality?
Yuki Tatsunami, Masato Taki
For the past ten years, CNN has reigned supreme in the world of computer vision, but recently, Transformer has been on the rise. However, the quadratic computational cost of self-a…