1 paper
Zan-Xia Jin, Heran Wu, Chun Yang +4
Text-based visual question answering (VQA) requires to read and understand text in an image to correctly answer a given question. However, most current methods simply add optical c…