5 papers
Towards One-to-Many Temporal Grounding
Qi Xu, Yue Tan, Shihao Chen +5
Temporal Grounding (TG) aims to localize video segments corresponding to a textual query. Prior research predominantly focuses on single-segment retrieval. Real-world scenarios, ho…
Watch, Remember, Reason: Human-View Video Understanding with MLLMs
Jiahao Meng, Yue Tan, Qi Xu +12
Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video…
EvoSpec: Evolving Speculative Decoding via Real-Time Vocabulary and Parameter Adaptation
Shuyu Zhang, Lingfeng Pan, Qicheng Wang +6
Speculative decoding accelerates Large Language Model inference through draft-then-verify generation, yet lightweight draft models face coupled efficiency and quality limitations:…
Open-o3-Video: Grounded Video Reasoning with Explicit Spatio-Temporal Evidence
Jiahao Meng, Xiangtai Li, Haochen Wang +8
Most video reasoning models only generate textual reasoning traces without indicating when and where key evidence appears. Recent models such as OpenAI-o3 have sparked wide interes…
CyberV: Cybernetics for Test-time Scaling in Video Understanding
Jiahao Meng, Shuyang Sun, Yue Tan +4
Current Multimodal Large Language Models (MLLMs) may struggle with understanding long or complex videos due to computational demands at test time, lack of robustness, and limited a…