5 papers
Seeing Straight: Document Orientation Detection for Efficient OCR
Suranjan Goswami, Abhinav Ravi, Raja Kolla +5
Despite significant advances in document understanding, determining the correct orientation of scanned or photographed documents remains a critical pre-processing step in the real…
Chitrakshara: A Large Multilingual Multimodal Dataset for Indian languages
Shaharukh Khan, Ali Faraz, Abhinav Ravi +6
Multimodal research has predominantly focused on single-image reasoning, with limited exploration of multi-image scenarios. Recent models have sought to enhance multi-image underst…
Designing Production-Scale OCR for India: Multilingual and Domain-Specific Systems
Ali Faraz, Raja Kolla, Ashish Kulkarni +1
Designing Optical Character Recognition (OCR) systems for India requires balancing linguistic diversity, document heterogeneity, and deployment constraints. In this paper, we study…
IndicVisionBench: Benchmarking Cultural and Multilingual Understanding in VLMs
Ali Faraz, Akash, Shaharukh Khan +6
Vision-language models (VLMs) have demonstrated impressive generalization across multimodal tasks, yet most evaluation benchmarks remain Western-centric, leaving open questions abo…
Chitrarth: Bridging Vision and Language for a Billion People
Shaharukh Khan, Ayush Tarun, Abhinav Ravi +7
Recent multimodal foundation models are primarily trained on English or high resource European language data, which hinders their applicability to other medium and low-resource lan…