2 citations · 2 across the 3 of their papers we have counts for
1 paper · 1 filter
Shicheng Yin, Kaixuan Yin, Yang Liu +2
The content-agnostic, fixed-grid tokenizers used by standard large-scale vision models like Vision Transformer (ViT) and Vision Mamba (Vim) represent a fundamental performance bott…