2 papers
cs.CV2026
Beyond OCR Accuracy: Text-Centric VQA Under Image Degradation with Modular and End-to-End
Ritali Vatsi, Rachapudi Jagadeesh, Shruti Singh Baghel +3
Text-centric Visual Question Answering (VQA) requires reading and reasoning over text embedded in images, a task made substantially harder when images suffer from real-world degrad…
cs.CV2025
Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
Shruti Singh Baghel, Yash Pratap Singh Rathore, Sushovan Jena +4
Large Vision-Language Models (VLMs) excel at understanding and generating video descriptions but their high memory, computation, and deployment demands hinder practical use particu…