3 citations · 5 across the 2 of their papers we have counts for
4 papers
Analyzing the Efficacy of an LLM-Only Approach for Image-based Document Question Answering
Nidhi Hegde, Sujoy Paul, Gagan Madan +1
Recent document question answering models consist of two key components: the vision encoder, which captures layout and visual elements in images, and a Large Language Model (LLM) t…
Is it an i or an l: Test-time Adaptation of Text Line Recognition Models
Debapriya Tula, Sujoy Paul, Gagan Madan +3
Recognizing text lines from images is a challenging problem, especially for handwritten documents due to large variations in writing styles. While text line recognition models are…
Weakly supervised information extraction from inscrutable handwritten document images
Sujoy Paul, Gagan Madan, Akankshya Mishra +3
State-of-the-art information extraction methods are limited by OCR errors. They work well for printed text in form-like documents, but unstructured, handwritten documents still rem…
A Study of Autoregressive Decoders for Multi-Tasking in Computer Vision
Lucas Beyer, Bo Wan, Gagan Madan +9
There has been a recent explosion of computer vision models which perform many tasks and are composed of an image encoder (usually a ViT) and an autoregressive decoder (usually a T…