733 citations · 899 across the 14 of their papers we have counts for
1 paper · 1 filter
Francesco Pinto, Nathalie Rauschmayr, Florian Tramèr +2
Vision-Language Models (VLMs) have made remarkable progress in document-based Visual Question Answering (i.e., responding to queries about the contents of an input document provide…