activity
20242026
collaborators

6 papers

cs.CV2026

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1

Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and e…

cs.CV2026

Don't Pause! Every prediction matters in a streaming video

Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener +2

Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause…

cs.CV2026

On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding

Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1

Multimodal Large Language Models (MLLMs) have advanced open-world action understanding and can be adapted as generative classifiers for closed-set settings by autoregressively gene…

cs.CV2025

Context-Enhanced Memory-Refined Transformer for Online Action Detection

Zhanzhong Pang, Fadime Sener, Angela Yao

Online Action Detection (OAD) detects actions in streaming videos using past observations. State-of-the-art OAD approaches model past observations and their interactions with an an…

cs.CV2025

Cost-Sensitive Learning for Long-Tailed Temporal Action Segmentation

Zhanzhong Pang, Fadime Sener, Shrinivas Ramasubramanian +1

Temporal action segmentation in untrimmed procedural videos aims to densely label frames into action classes. These videos inherently exhibit long-tailed distributions, where actio…

cs.CV2024

Long-Tail Temporal Action Segmentation with Group-wise Temporal Logit Adjustment

Zhanzhong Pang, Fadime Sener, Shrinivas Ramasubramanian +1

Procedural activity videos often exhibit a long-tailed action distribution due to varying action frequencies and durations. However, state-of-the-art temporal action segmentation m…