1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Yi Zhu, Yanpeng Zhou, Chunwei Wang +4
Currently, vision encoder models like Vision Transformers (ViTs) typically excel at image recognition tasks but cannot simultaneously support text recognition like human visual rec…