5 citations · 5 across the 25 of their papers we have counts for
30 papers
TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models
Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh +1
Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in safety mechanisms. Existing concept erasur…
IP Protection in the Era of Visual Generative AI: A Survey
Zhuan Shi, Shunchang Liu, Alireza Dehghanpour Farashah +8
The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, reproduction, extraction, misuse, and re…
Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems
Marylou Fauchard, Florian Carichon, Margarida Carvalho +1
Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic dec…
Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
Ali khalil, Aly M. Kassem, Mohamed Abdelrazek +3
We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using…
IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts
Ayana Hussain, Soumya Sharma, Golnoosh Farnadi +3
Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evalua…
Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4
Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show th…