3 papers
cs.CV2026
Attend to what I say: Highlighting relevant content on slides
Megha Mariam K M, C. V. Jawahar
Imagine sitting in a presentation, trying to follow the speaker while simultaneously scanning the slides for relevant information. While the entire slide is visible, identifying th…
cs.CV2025
HW-MLVQA: Elucidating Multilingual Handwritten Document Understanding with a Comprehensive VQA Benchmark
Aniket Pal, Ajoy Mondal, Minesh Mathew +1
The proliferation of MultiLingual Visual Question Answering (MLVQA) benchmarks augments the capabilities of large language models (LLMs) and multi-modal LLMs, thereby enabling them…
cs.CV2025
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
Soumya Shamarao Jahagirdar, Jayasree Saha, C V Jawahar
Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging whe…