5 papers · 1 filter
Decouple and Cache: KV Cache Construction for Streaming Video Understanding
Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1
Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and e…
Don't Pause! Every prediction matters in a streaming video
Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener +2
Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause…
On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding
Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1
Multimodal Large Language Models (MLLMs) have advanced open-world action understanding and can be adapted as generative classifiers for closed-set settings by autoregressively gene…
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
Dibyadip Chatterjee, Edoardo Remelli, Yale Song +9
We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - ve…
On the Utility of 3D Hand Poses for Action Recognition
Md Salman Shamil, Dibyadip Chatterjee, Fadime Sener +2
3D hand pose is an underexplored modality for action recognition. Poses are compact yet informative and can greatly benefit applications with limited compute budgets. However, pose…