5 papers · 1 filter
How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA
Navya Gupta, Bingjie Xu, Avinash Anand +2
Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attr…
Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning
Janak Kapuriya, Anwar Shaikh, Arnav Goel +8
In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLM…
Keystroke Dynamics Against Academic Dishonesty in the Age of LLMs
Debnath Kundu, Atharva Mehta, Rajesh Kumar +4
The transition to online examinations and assignments raises significant concerns about academic integrity. Traditional plagiarism detection systems often struggle to identify inst…
TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content
Avinash Anand, Raj Jaiswal, Pijush Bhuyan +5
The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable…
RanLayNet: A Dataset for Document Layout Detection used for Domain Adaptation and Generalization
Avinash Anand, Raj Jaiswal, Mohit Gupta +7
Large ground-truth datasets and recent advances in deep learning techniques have been useful for layout detection. However, because of the restricted layout diversity of these data…