2 citations · 2 across the 3 of their papers we have counts for
4 papers · 1 filter
The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models
Adeel Yousaf, Soumik Ghosh, James Beetham +2
Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver hig…
SafeR-CLIP: Mitigating NSFW Content in Vision-Language Models While Preserving Pre-Trained Knowledge
Adeel Yousaf, Joseph Fioresi, James Beetham +2
Improving the safety of vision-language models like CLIP via fine-tuning often comes at a steep price, causing significant drops in their generalization performance. We find this t…
Videoprompter: an ensemble of foundational models for zero-shot video understanding
Adeel Yousaf, Muzammal Naseer, Salman Khan +2
Vision-language models (VLMs) classify the query video by calculating a similarity score between the visual features and text-based class label representations. Recently, large lan…
EventTransAct: A video transformer-based framework for Event-camera based action recognition
Tristan de Blegiers, Ishan Rajendrakumar Dave, Adeel Yousaf +1
Recognizing and comprehending human actions and gestures is a crucial perception requirement for robots to interact with humans and carry out tasks in diverse domains, including se…