3 papers
cs.CV2026
TrustCLIP: Learning Private Visual Features via Adversarial Reconstruction
Nikos Athanasiou, Ilya A. Petrov, Angela Yao +8
Vision and vision-language models rely on high-level visual representations that are increasingly used across recognition, retrieval, and multimodal reasoning pipelines. However, r…
cs.CV2025
Memory-efficient Streaming VideoLLMs for Real-time Procedural Video Understanding
Dibyadip Chatterjee, Edoardo Remelli, Yale Song +9
We introduce ProVideLLM, an end-to-end framework for real-time procedural video understanding. ProVideLLM integrates a multimodal cache configured to store two types of tokens - ve…
cs.CV2024
DiffH2O: Diffusion-Based Synthesis of Hand-Object Interactions from Textual Descriptions
Sammy Christen, Shreyas Hampali, Fadime Sener +5
Generating natural hand-object interactions in 3D is challenging as the resulting hand and object motions are expected to be physically plausible and semantically meaningful. Furth…