16 citations · 30 across the 8 of their papers we have counts for
1 paper · 1 filter
Yongxin Zhu, Zhen Liu, Yukang Liang +4
In this paper, we propose a novel multi-modal framework for Scene Text Visual Question Answering (STVQA), which requires models to read scene text in images for question answering.…