collaborators

12 papers

cs.CV2026

TEA: Text Encoder Alignment for Robust Concept Erasure in Text-to-Image Models

Alireza Dehghanpour Farashah, Zhuan Shi, Negar Rostamzadeh +1

Text-to-image diffusion models can be misused to generate harmful content through adversarial or paraphrased prompts that bypass built-in safety mechanisms. Existing concept erasur…

cs.CR2026

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

Ali khalil, Aly M. Kassem, Mohamed Abdelrazek +3

We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using…

cs.AI2026

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4

Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show th…

cs.AI2026

A Unified Framework to Quantify Cultural Intelligence of AI

Sunipa Dev, Vinodkumar Prabhakaran, Rutledge Chin Feman +16

As generative AI technologies are increasingly being launched across the globe, assessing their competence to operate in different cultural contexts is exigently becoming a priorit…

cs.LG2026

Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes

Aly Kassem, Thomas Jiralerspong, Negar Rostamzadeh +1

Model diffing methods aim to identify how fine-tuning changes a model's internal representations. Crosscoders approach this by learning shared dictionaries of interpretable latent…

cs.CL2026

Multilingual Amnesia: On the Transferability of Unlearning in Multilingual LLMs

Alireza Dehghanpour Farashah, Aditi Khandelwal, Marylou Fauchard +3

As multilingual large language models become more widely used, ensuring their safety and fairness across diverse linguistic contexts presents unique challenges. While existing rese…