1 paper
Kjetil Indrehus, Adrian Duric, Changkyu Choi +1
Document Visual Question Answering (DocVQA) requires vision-language models to reason not only about what information in a document is relevant to a question, but also where the an…