10 citations · 33 across the 6 of their papers we have counts for
Showing 2022Show all
2 papers · 1 filter
cs.CV2022★ 5 cited
Masked Vision-Language Transformers for Scene Text Recognition
Jie Wu, Ying Peng, Shengming Zhang +2
Scene text recognition (STR) enables computers to recognize and read the text in various real-world scenes. Recent STR models benefit from taking linguistic information in addition…
cs.CV2022
ScalableViT: Rethinking the Context-oriented Generalization of Vision Transformer
Rui Yang, Hailong Ma, Jie Wu +4
The vanilla self-attention mechanism inherently relies on pre-defined and steadfast computational dimensions. Such inflexibility restricts it from possessing context-oriented gener…