4 papers
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
Jinlong Yang, Wenhao Zhang, Kuanwei Lin +1
Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example rega…
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2
Long-video reasoning is fundamentally constrained by how models acquire and utilize visual evidence. Existing tool-augmented video frameworks often interleave temporal grounding an…
TIR-Flow: Active Video Search and Reasoning with Frozen VLMs
Hongbo Jin, Siyi Xie, Jiayu Ding +2
While Large Video-Language Models (Video-LLMs) have achieved remarkable progress in perception, their reasoning capabilities remain a bottleneck. Existing solutions typically resor…
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
Hongbo Jin, Kuanwei Lin, Wenhao Zhang +2
Reinforcement Learning (RL) is crucial for empowering VideoLLMs with complex spatiotemporal reasoning. However, current RL paradigms predominantly rely on random data shuffling or…