3 papers
cs.CV2026
AdaVDR: Adaptive Tool Use and Reflection for Video Deep Research
Xintong Zhang, Xiaomeng Fan, Shilin Yan +7
Video deep research answers complex questions by jointly understanding video content and retrieving external knowledge from the open Web. However, diverse questions and videos requ…
cs.CV2026
Bridging Modality Disconnect in Self-Reflection via Closed-Loop Visually Grounded Verification
Haoyu Zhang, Yuwei Wu, Pengxiang Li +6
In the era of Vision-Language Models (VLMs), enhancing multimodal reasoning capabilities remains a critical challenge, particularly in handling ambiguous or complex visual inputs,…
cs.RO2025
Long-Horizon Visual Imitation Learning via Plan and Code Reflection
Quan Chen, Chenrui Shi, Qi Chen +6
Learning from long-horizon demonstrations with complex action sequences presents significant challenges for visual imitation learning, particularly in understanding temporal relati…