Showing cs.CVShow all
2 papers · 1 filter
cs.CV2026
Evaluating Self-Correcting Vision Agents Through Quantitative and Qualitative Metrics
Aradhya Dixit
Recent progress in multimodal foundation models has enabled Vision-Language Agents (VLAs) to decompose complex visual tasks into executable tool-based plans. While recent benchmark…
cs.CV2026
Semantic Event Graphs for Long-Form Video Question Answering
Aradhya Dixit, Tianxi Liang
Long-form video question answering remains challenging for modern vision-language models, which struggle to reason over hour-scale footage without exceeding practical token and com…