2 citations · 2 across the 3 of their papers we have counts for
3 papers
cs.CV2023
Analyzing the Efficacy of an LLM-Only Approach for Image-based Document Question Answering
Nidhi Hegde, Sujoy Paul, Gagan Madan +1
Recent document question answering models consist of two key components: the vision encoder, which captures layout and visual elements in images, and a Large Language Model (LLM) t…
cs.CV2023
Is it an i or an l: Test-time Adaptation of Text Line Recognition Models
Debapriya Tula, Sujoy Paul, Gagan Madan +3
Recognizing text lines from images is a challenging problem, especially for handwritten documents due to large variations in writing styles. While text line recognition models are…
cs.CV2023★ 2 cited
Weakly supervised information extraction from inscrutable handwritten document images
Sujoy Paul, Gagan Madan, Akankshya Mishra +3
State-of-the-art information extraction methods are limited by OCR errors. They work well for printed text in form-like documents, but unstructured, handwritten documents still rem…