2 papers
cs.CV2026
CADER: Confidence-Aware Dynamic Evidence Reasoning for Long-Video Understanding
Jinlong Yang, Wenhao Zhang, Kuanwei Lin +1
Long-video understanding increasingly relies on large vision-language models and tool-augmented reasoning, but most systems apply the same inference procedure to every example rega…
cs.CV2025
VideoMem: Enhancing Ultra-Long Video Understanding via Adaptive Memory Management
Hongbo Jin, Qingyuan Wang, Wenhao Zhang +2
Ultra long video understanding remains an open challenge, as existing vision language models (VLMs) falter on such content due to limited context length and inefficient long term m…