1 citations · 2 across the 7 of their papers we have counts for
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2025
Linguistics-aware Masked Image Modeling for Self-supervised Scene Text Recognition
Yifei Zhang, Chang Liu, Jin Wei +4
Text images are unique in their dual nature, encompassing both visual and linguistic information. The visual component encompasses structural and appearance-based features, while t…
cs.CV2024★ 1 cited
Track the Answer: Extending TextVQA from Image to Video with Spatio-Temporal Clues
Yan Zhang, Gangyan Zeng, Huawen Shen +3
Video text-based visual question answering (Video TextVQA) is a practical task that aims to answer questions by jointly reasoning textual and visual information in a given video. I…
cs.CV2023★ 1 cited
Separate and Locate: Rethink the Text in Text-based Visual Question Answering
Chengyang Fang, Jiangnan Li, Liang Li +2
Text-based Visual Question Answering (TextVQA) aims at answering questions about the text in images. Most works in this field focus on designing network structures or pre-training…