62 citations · 70 across the 6 of their papers we have counts for
4 papers
Multiple-Question Multiple-Answer Text-VQA
Peng Tang, Srikar Appalaraju, R. Manmatha +2
We present Multiple-Question Multiple-Answer (MQMA), a novel approach to do text-VQA in encoder-decoder transformer models. The text-VQA task requires a model to answer a question…
SimCon Loss with Multiple Views for Text Supervised Semantic Segmentation
Yash Patel, Yusheng Xie, Yi Zhu +2
Learning to segment images purely by relying on the image-text alignment from web data can lead to sub-optimal performance due to noise in the data. The noise comes from the sample…
AIM: Adapting Image Models for Efficient Video Action Recognition
Taojiannan Yang, Yi Zhu, Yusheng Xie +3
Recent vision transformer based video models mostly follow the ``image pre-training then finetuning" paradigm and have achieved great success on multiple video benchmarks. However,…
LaTr: Layout-Aware Transformer for Scene-Text VQA
Ali Furkan Biten, Ron Litman, Yusheng Xie +2
We propose a novel multimodal architecture for Scene Text Visual Question Answering (STVQA), named Layout-Aware Transformer (LaTr). The task of STVQA requires models to reason over…