5 citations · 10 across the 11 of their papers we have counts for
1 paper · 2 filters
Xuejing Liu, Wei Tang, Xinzhe Ni +4
Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this…