5 papers · 1 filter
StreamForest: Efficient Online Video Understanding with Persistent Event Memory
Xiangyu Zeng, Kefan Qiu, Qingyu Zhang +9
Multimodal Large Language Models (MLLMs) have recently achieved remarkable progress in video understanding. However, their effectiveness in real-time streaming scenarios remains li…
Online Video Understanding: OVBench and VideoChat-Online
Zhenpeng Huang, Xinhao Li, Jiaqi Li +7
Multimodal Large Language Models (MLLMs) have significantly progressed in offline video understanding. However, applying these models to real-world scenarios, such as autonomous dr…
Learning Human Skill Generators at Key-Step Levels
Yilu Wu, Chenhui Zhu, Shuai Wang +4
We are committed to learning human skill generators at key-step levels. The generation of skills is a challenging endeavor, but its successful implementation could greatly facilita…
VideoEval: Comprehensive Benchmark Suite for Low-Cost Evaluation of Video Foundation Model
Xinhao Li, Zhenpeng Huang, Jing Wang +2
With the growth of high-quality data and advancement in visual pre-training paradigms, Video Foundation Models (VFMs) have made significant progress recently, demonstrating their r…
Open-Event Procedure Planning in Instructional Videos
Yilu Wu, Hanlin Wang, Jing Wang +1
Given the current visual observations, the traditional procedure planning task in instructional videos requires a model to generate goal-directed plans within a given action space.…