dataset creation 1explainable AI 1spatio-temporal grounding 1video large language models 1video question answering 1
From the 1 of 4 linked papers with an AI index.
Showing cs.CVShow all
3 papers · 1 filter
cs.CV2026
Evidence-Backed Video Question Answering
Shijie Wang, Honglu Zhou, Ziyang Wang +5
The paper introduces Evidence-Backed Video Question Answering (E-VQA), a task where models must provide both a textual answer and precise spatio‑temporal visual evidence (temporal…
cs.CV2026
Towards Robust Sequential Decomposition for Complex Image Editing
Zilai Zeng, Mingdeng Cao, Zijie Li +5
Recent advances in visual generative models have enabled high-fidelity image editing guided by human instructions. However, these models often struggle with complex instructions in…
cs.CV2024
Dense Video Object Captioning from Disjoint Supervision
Xingyi Zhou, Anurag Arnab, Chen Sun +1
We propose a new task and model for dense video object captioning -- detecting, tracking and captioning trajectories of objects in a video. This task unifies spatial and temporal l…