12 citations · 23 across the 5 of their papers we have counts for
10 papers · 1 filter
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
Aniket Pal, Ajoy Mondal, Minesh Mathew +1
The proliferation of MultiLingual Visual Question Answering (MLVQA) benchmarks augments the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them…
Asking questions on handwritten document collections
Minesh Mathew, Lluis Gomez, Dimosthenis Karatzas +1
This work addresses the problem of Question Answering (QA) on handwritten document collections. Unlike typical QA and Visual Question Answering (VQA) formulations where the answer…
Benchmarking Scene Text Recognition in Devanagari, Telugu and Malayalam
Minesh Mathew, Mohit Jain, CV Jawahar
Inspired by the success of Deep Learning based approaches to English scene text recognition, we pose and benchmark scene text recognition for three Indic scripts - Devanagari, Telu…
MMBERT: Multimodal BERT Pretraining for Improved Medical VQA
Yash Khare, Viraj Bagal, Minesh Mathew +3
Images in the medical domain are fundamentally different from the general domain images. Consequently, it is infeasible to directly employ general domain Visual Question Answering…
InfographicVQA
Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3
Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic und…
DocVQA: A Dataset for VQA on Document Images
Minesh Mathew, Dimosthenis Karatzas, C. V. Jawahar
We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed…