activity
20242026
collaborators
Showing cs.CVShow all

5 papers · 1 filter

cs.CV2026

How Do VLMs Fail? Vision-Operation Misalignment in Compositional VQA

Navya Gupta, Bingjie Xu, Avinash Anand +2

Compositional visual question answering requires Vision-Language Models (VLMs) to execute multiple reasoning operations like object selection, spatial relation resolution, and attr…

cs.CV2025

Enhancing Scientific Visual Question Answering via Vision-Caption aware Supervised Fine-Tuning

Janak Kapuriya, Anwar Shaikh, Arnav Goel +8

In this study, we introduce Vision-Caption aware Supervised FineTuning (VCASFT), a novel learning paradigm designed to enhance the performance of smaller Vision Language Models(VLM…

cs.CV2024

Keystroke Dynamics Against Academic Dishonesty in the Age of LLMs

Debnath Kundu, Atharva Mehta, Rajesh Kumar +4

The transition to online examinations and assignments raises significant concerns about academic integrity. Traditional plagiarism detection systems often struggle to identify inst…

cs.CV2024

TC-OCR: TableCraft OCR for Efficient Detection & Recognition of Table Structure & Content

Avinash Anand, Raj Jaiswal, Pijush Bhuyan +5

The automatic recognition of tabular data in document images presents a significant challenge due to the diverse range of table styles and complex structures. Tables offer valuable…

cs.CV2024

RanLayNet: A Dataset for Document Layout Detection used for Domain Adaptation and Generalization

Avinash Anand, Raj Jaiswal, Mohit Gupta +7

Large ground-truth datasets and recent advances in deep learning techniques have been useful for layout detection. However, because of the restricted layout diversity of these data…