activity
20172025
most citedAsking questions on handwritten document collections

12 citations · 23 across the 5 of their papers we have counts for

collaborators
Showing cs.CVShow all

10 papers · 1 filter

cs.CV2025

HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark

Aniket Pal, Ajoy Mondal, Minesh Mathew +1

The proliferation of MultiLingual Visual Question Answering (MLVQA) benchmarks augments the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them…

cs.CV202112 cited

Asking questions on handwritten document collections

Minesh Mathew, Lluis Gomez, Dimosthenis Karatzas +1

This work addresses the problem of Question Answering (QA) on handwritten document collections. Unlike typical QA and Visual Question Answering (VQA) formulations where the answer…

cs.CV2021

Benchmarking Scene Text Recognition in Devanagari, Telugu and Malayalam

Minesh Mathew, Mohit Jain, CV Jawahar

Inspired by the success of Deep Learning based approaches to English scene text recognition, we pose and benchmark scene text recognition for three Indic scripts - Devanagari, Telu…

cs.CV2021

MMBERT: Multimodal BERT Pretraining for Improved Medical VQA

Yash Khare, Viraj Bagal, Minesh Mathew +3

Images in the medical domain are fundamentally different from the general domain images. Consequently, it is infeasible to directly employ general domain Visual Question Answering…

cs.CV2021

InfographicVQA

Minesh Mathew, Viraj Bagal, Rubèn Pérez Tito +3

Infographics are documents designed to effectively communicate information using a combination of textual, graphical and visual elements. In this work, we explore the automatic und…

cs.CV2020

DocVQA: A Dataset for VQA on Document Images

Minesh Mathew, Dimosthenis Karatzas, C. V. Jawahar

We present a new dataset for Visual Question Answering (VQA) on document images called DocVQA. The dataset consists of 50,000 questions defined on 12,000+ document images. Detailed…