4 papers
DRISHTIKON: Visual Grounding at Multiple Granularities in Documents
Badri Vishal Kasuba, Parag Chaudhuri, Ganesh Ramakrishnan
Visual grounding in text-rich document images is a critical yet underexplored challenge for Document Intelligence and Visual Question Answering (VQA) systems. We present DRISHTIKON…
SPRINT: Script-agnostic Structure Recognition in Tables
Dhruv Kudale, Badri Vishal Kasuba, Venkatapathy Subramanian +2
Table Structure Recognition (TSR) is vital for various downstream tasks like information retrieval, table reconstruction, and document understanding. While most state-of-the-art (S…
PLATTER: A Page-Level Handwritten Text Recognition System for Indic Scripts
Badri Vishal Kasuba, Dhruv Kudale, Venkatapathy Subramanian +2
In recent years, the field of Handwritten Text Recognition (HTR) has seen the emergence of various new models, each claiming to perform competitively better than the other in speci…
StructFormer: Document Structure-based Masked Attention and its Impact on Language Model Pre-Training
Kaustubh Ponkshe, Venkatapathy Subramanian, Natwar Modani +1
Most state-of-the-art techniques for Language Models (LMs) today rely on transformer-based architectures and their ubiquitous attention mechanism. However, the exponential growth i…