3 citations · 4 across the 2 of their papers we have counts for
1 paper · 1 filter
Yixuan Qiao, Hao Chen, Jun Wang +7
TextVQA requires models to read and reason about text in images to answer questions about them. Specifically, models need to incorporate a new modality of text present in the image…