collaborators

5 papers

cs.LG2026

ProactiveBench: Can Streaming Video Models Really Interact Like Humans?

Kaixuan Du, Xin Wan, YuKun Wang +5

Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query…

cs.CV2026

Putting Captions to the Test: Evaluating Video Caption Quality through Multiple-Choice Question Answering

Zizhen Wang, Bo Feng, Zhengfeng Lai +5

Evaluating video captioning remains a critical challenge for Visual Large Language Models (VLLMs). Existing metrics primarily rely on matching generated text against ground-truth r…

cs.CV2026

VSAS-Bench: Real-Time Evaluation of Visual Streaming Assistant Models

Pavan Kumar Anasosalu Vasu, Cem Koc, Fartash Faghri +6

Streaming vision-language models (VLMs) continuously generate responses given an instruction prompt and an online stream of input frames. This is a core mechanism for real-time vis…

cs.CV2025

Breaking Down Video LLM Benchmarks: Knowledge, Spatial Perception, or True Temporal Understanding?

Bo Feng, Zhengfeng Lai, Shiyu Li +4

Existing video understanding benchmarks often conflate knowledge-based and purely image-based questions, rather than clearly isolating a model's temporal reasoning ability, which i…

cs.CV2025

StreamBridge: Turning Your Offline Video Large Language Model into a Proactive Streaming Assistant

Haibo Wang, Bo Feng, Zhengfeng Lai +6

We present StreamBridge, a simple yet effective framework that seamlessly transforms offline Video-LLMs into streaming-capable models. It addresses two fundamental challenges in ad…