1 citations · 1 across the 1 of their papers we have counts for
1 paper · 1 filter
Xiaopeng Lu, Zhen Fan, Yansen Wang +2
As an important task in multimodal context understanding, Text-VQA (Visual Question Answering) aims at question answering through reading text information in images. It differentia…