12 citations · 12 across the 1 of their papers we have counts for
1 paper · 1 filter
Chongjian Ge, Xiaohan Ding, Zhan Tong +4
Vision Transformers (ViTs) have been shown to enhance visual recognition through modeling long-range dependencies with multi-head self-attention (MHSA), which is typically formulat…