From the 1 of 5 linked papers with an AI index.
5 papers
EFlow: Learning Evidence Flow for Long-Video Reasoning with Adaptive Reflection
Wenhao Zhang, Kuanwei Lin, Xuyi Yang +2
The paper introduces EFlow, a framework that first retrieves visual evidence from long videos before reasoning, using separate chain‑of‑thought modules for temporal grounding and a…
StereoVLA: Enhancing Vision-Language-Action Models with Stereo Vision
Shengliang Deng, Mi Yan, Yixin Zheng +7
While Vision-Language-Action (VLA) models excel in generalist manipulation, they often lack fine-grained spatial awareness and show limited viewpoint robustness. This limitation la…
VideoCuRL: Video Curriculum Reinforcement Learning with Orthogonal Difficulty Decomposition
Hongbo Jin, Kuanwei Lin, Wenhao Zhang +2
Reinforcement Learning (RL) is crucial for empowering VideoLLMs with complex spatiotemporal reasoning. However, current RL paradigms predominantly rely on random data shuffling or…
Global Convergence of Policy Gradient for Entropy Regularized Linear-Quadratic Control with Multiplicative Noise
Gabriel Diaz, Lucky Li, Wenhao Zhang
Reinforcement Learning (RL) has emerged as a powerful framework for sequential decision-making in dynamic environments, particularly when system parameters are unknown. This paper…
Enhancing Long Video Question Answering with Scene-Localized Frame Grouping
Xuyi Yang, Wenhao Zhang, Hongbo Jin +5
Current Multimodal Large Language Models (MLLMs) often perform poorly in long video understanding, primarily due to resource limitations that prevent them from processing all video…