5 citations · 7 across the 5 of their papers we have counts for
3 papers · 1 filter
CRAG-MM-Diagnostics: Enabling Stage-Wise Analysis of Knowledge-Intensive VQA
Hanseok Oh, Parishad BehnamGhader, Benno Krojer +4
Knowledge-Intensive Visual Question Answering (KI-VQA) benchmarks evaluate Vision-Language Models (VLMs) as multimodal knowledge assistants by requiring external information beyond…
Zero-Shot Dense Video Captioning by Jointly Optimizing Text and Moment
Yongrae Jo, Seongyun Lee, Aiden SJ Lee +3
Dense video captioning, a task of localizing meaningful moments and generating relevant captions for videos, often requires a large, expensive corpus of annotated video segments pa…
ViSeRet: A simple yet effective approach to moment retrieval via fine-grained video segmentation
Aiden Seungjoon Lee, Hanseok Oh, Minjoon Seo
Video-text retrieval has many real-world applications such as media analytics, surveillance, and robotics. This paper presents the 1st place solution to the video retrieval track o…