activity
20242026
collaborators

10 papers

cs.CV2026

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

Nikos Athanasiou, Ilya A. Petrov, Angela Yao +8

Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal reasoning pipelines. However, r…

cs.CV2026

Decouple and Cache: KV Cache Construction for Streaming Video Understanding

Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1

Streaming video understanding requires processing unbounded video streams with limited memory and computation, posing two key challenges. First, continuously constructing new and e…

cs.CV2026

Don't Pause! Every prediction matters in a streaming video

Dibyadip Chatterjee, Zhanzhong Pang, Fadime Sener +2

Streaming video models should respond the moment an event unfolds, not after the moment has passed. Yet existing online VideoQA benchmarks remain largely retrospective. They pause…

cs.CV2026

On Discriminative vs. Generative classifiers: Rethinking MLLMs for Action Understanding

Zhanzhong Pang, Dibyadip Chatterjee, Fadime Sener +1

Multimodal Large Language Models (MLLMs) have advanced open-world action understanding and can be adapted as generative classifiers for closed-set settings by autoregressively gene…

cs.CV2026

PALM: A Dataset and Baseline for Learning Multi-subject Hand Prior

Zicong Fan, Edoardo Remelli, David Dimond +5

The ability to grasp objects, signal with gestures, and share emotion through touch all stem from the unique capabilities of human hands. Yet creating high-quality personalized han…

cs.CV2025

SneakPeek: Future-Guided Instructional Streaming Video Generation

Cheeun Hong, German Barquero, Fadime Sener +6

Instructional video generation is an emerging task that aims to synthesize coherent demonstrations of procedural activities from textual descriptions. Such capability has broad imp…