4 citations · 6 across the 4 of their papers we have counts for
6 papers
Analyzing the Efficacy of an LLM-Only Approach for Image-based Document Question Answering
Nidhi Hegde, Sujoy Paul, Gagan Madan +1
Recent document question answering models consist of two key components: the vision encoder, which captures layout and visual elements in images, and a Large Language Model (LLM) t…
Is it an i or an l: Test-time Adaptation of Text Line Recognition Models
Debapriya Tula, Sujoy Paul, Gagan Madan +3
Recognizing text lines from images is a challenging problem, especially for handwritten documents due to large variations in writing styles. While text line recognition models are…
Weakly supervised information extraction from inscrutable handwritten document images
Sujoy Paul, Gagan Madan, Akankshya Mishra +3
State-of-the-art information extraction methods are limited by OCR errors. They work well for printed text in form-like documents, but unstructured, handwritten documents still rem…
GAN-MPC: Training Model Predictive Controllers with Parameterized Cost Functions using Demonstrations from Non-identical Experts
Returaj Burnwal, Anirban Santara, Nirav P. Bhatt +2
Model predictive control (MPC) is a popular approach for trajectory optimization in practical robotics applications. MPC policies can optimize trajectory parameters under kinodynam…
Two Central limit theorems in Diophantine approximation
Gaurav Aggarwal, Anish Ghosh
We prove central limit theorems for Diophantine approximations with congruence conditions and for inhomogeneous Diophantine approximations following the approach of Björklund and G…
The Beauty of Capturing Faces: Rating the Quality of Digital Portraits
Miriam Redi, Nikhil Rasiwasia, Gaurav Aggarwal +1
Digital portrait photographs are everywhere, and while the number of face pictures keeps growing, not much work has been done to on automatic portrait beauty assessment. In this pa…