1 paper
Xuejing Liu, Wei Tang, Xinzhe Ni +4
Text-rich VQA, namely Visual Question Answering based on text recognition in the images, is a cross-modal task that requires both image comprehension and text recognition. In this…