2 citations · 2 across the 1 of their papers we have counts for
1 paper
Ali Furkan Biten, Ron Litman, Yusheng Xie +2
We propose a novel multimodal architecture for Scene Text Visual Question Answering (STVQA), named Layout-Aware Transformer (LaTr). The task of STVQA requires models to reason over…