8 citations · 16 across the 6 of their papers we have counts for
1 paper · 1 filter
Xiaopeng Lu, Zhen Fan, Yansen Wang +2
As an important task in multimodal context understanding, Text-VQA (Visual Question Answering) aims at question answering through reading text information in images. It differentia…