6 papers
Prompt2LVideos: Exploring Prompts for Understanding Long-Form Multimodal Videos
Soumya Shamarao Jahagirdar, Jayasree Saha, C V Jawahar
Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging whe…
Understanding Video Scenes through Text: Insights from Text-based Video Question Answering
Soumya Jahagirdar, Minesh Mathew, Dimosthenis Karatzas +1
Researchers have extensively studied the field of vision and language, discovering that both visual and textual content is crucial for understanding scenes effectively. Particularl…
Making the V in Text-VQA Matter
Shamanthak Hegde, Soumya Jahagirdar, Shankar Gangisetty
Text-based VQA aims at answering questions by reading the text present in the images. It requires a large amount of scene-text relationship understanding compared to the VQA task.…
Weakly Supervised Visual Question Answer Generation
Charani Alampalle, Shamanthak Hegde, Soumya Jahagirdar +1
Growing interest in conversational agents promote twoway human-computer communications involving asking and answering visual questions have become an active area of research in AI.…
Look, Read and Ask: Learning to Ask Questions by Reading Text in Images
Soumya Jahagirdar, Shankar Gangisetty, Anand Mishra
We present a novel problem of text-based visual question generation or TextVQG in short. Given the recent growing interest of the document image analysis community in combining tex…
Watching the News: Towards VideoQA Models that can Read
Soumya Jahagirdar, Minesh Mathew, Dimosthenis Karatzas +1
Video Question Answering methods focus on commonsense reasoning and visual cognition of objects or persons and their interactions over time. Current VideoQA approaches ignore the t…