1 paper
Zhixuan Shen, Haonan Luo, Sijia Li +1
Scene-Text Visual Question Answering (ST-VQA) aims to understand scene text in images and answer questions related to the text content. Most existing methods heavily rely on the ac…