collaborators

5 papers

cs.CV2026

TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction

Nikos Athanasiou, Ilya A. Petrov, Angela Yao +7

Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal reasoning pipelines. However, r…

cs.CV2025

Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding

Dibyadip Chatterjee, Edoardo Remelli, Yale Song +9

We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - ve…

cs.CV2024

X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization

Anna Kukleva, Fadime Sener, Edoardo Remelli +4

Lately, there has been growing interest in adapting vision-language models (VLMs) to image and third-person video classification due to their success in zero-shot recognition. Howe…

cs.CV2024

DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions

Sammy Christen, Shreyas Hampali, Fadime Sener +5

Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful. Furth…

cs.CV2024

On the Utility of 3D Hand Poses for Action Recognition

Md Salman Shamil, Dibyadip Chatterjee, Fadime Sener +2

3D hand pose is an underexplored modality for action recognition. Poses are compact yet informative and can greatly benefit applications with limited compute budgets. However, pose…