3 papers
cs.CV2025
Video Spatial Reasoning with Object-Centric 3D Rollout
Haoran Tang, Meng Cao, Ruyang Liu +4
Recent advances in Multi-modal Large Language Models (MLLMs) have showcased remarkable capabilities in vision-language understanding. However, enabling robust video spatial reasoni…
cs.CV2025
Flow4Agent: Long-form Video Understanding via Motion Prior from Optical Flow
Ruyang Liu, Shangkun Sun, Haoran Tang +2
Long-form video understanding has always been a challenging problem due to the significant redundancy in both temporal and spatial contents. This challenge is further exacerbated b…
cs.CV2025
VISTA: Mitigating Semantic Inertia in Video-LLMs via Training-Free Dynamic Chain-of-Thought Routing
Hongbo Jin, Jiayu Ding, Siyi Xie +2
Recent advancements in Large Language Models have successfully transitioned towards System 2 reasoning, yet applying these paradigms to video understanding remains challenging. Whi…