activity
20242026
most citedThe Coming Crisis of Multi-Agent Misalignment: AI Alignment Must Be a Dynamic and Social Process

5 citations · 5 across the 25 of their papers we have counts for

collaborators

30 papers

cs.CV2026

TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models

Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh +1

Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in safety mechanisms. Existing concept erasur…

cs.CV2026

IP Protection in the Era of Visual Generative AI: A Survey

Zhuan Shi, Shunchang Liu, Alireza Dehghanpour Farashah +8

The rapid evolution of visual generative AI has introduced a wide range of intellectual property risks, spanning the unauthorized learning, reproduction, extraction, misuse, and re…

cs.AI2026

Even More Deception: Objective Misalignment in Mixed-Motive LLM Multi-Agent Systems

Marylou Fauchard, Florian Carichon, Margarida Carvalho +1

Large Language Models (LLMs)-powered multi-agent systems are increasingly deployed in mixed-motive environments, where agents operate under asymmetric information and strategic dec…

cs.CR2026

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

Ali khalil, Aly M. Kassem, Mohamed Abdelrazek +3

We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using…

cs.CR2026

IDP-Bench: Benchmarking ability of LLMs to protect personal information in interdependent privacy contexts

Ayana Hussain, Soumya Sharma, Golnoosh Farnadi +3

Large language models (LLMs) are becoming widely deployed as personal AI assistants with access to sensitive user data, making privacy a major challenge for their design and evalua…

cs.AI2026

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4

Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show th…