1 paper · 1 filter
Kaixuan Du, Xin Wan, YuKun Wang +5
Streaming video understanding requires models to process continuous multimodal input while maintaining temporal context. Existing evaluations are predominantly reactive: they query…