collaborators

6 papers

cs.CR2026

Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior

Ali khalil, Aly M. Kassem, Mohamed Abdelrazek +3

We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using…

cs.AI2026

Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs

Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4

Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show th…

cs.LG2026

Delta-Crosscoder: Robust Crosscoder Model Diffing in Narrow Fine-Tuning Regimes

Aly Kassem, Thomas Jiralerspong, Negar Rostamzadeh +1

Model diffing methods aim to identify how fine-tuning changes a model's internal representations. Crosscoders approach this by learning shared dictionaries of interpretable latent…

cs.CL2025

Reviving Your MNEME: Predicting The Side Effects of LLM Unlearning and Fine-Tuning via Sparse Model Diffing

Aly M. Kassem, Zhuan Shi, Negar Rostamzadeh +1

Large language models (LLMs) are frequently fine-tuned or unlearned to adapt to new tasks or eliminate undesirable behaviors. While existing evaluation methods assess performance a…

cs.CL2025

How Robust Are Router-LLMs? Analysis of the Fragility of LLM Routing Capabilities

Aly M. Kassem, Bernhard Schölkopf, Zhijing Jin

Large language model (LLM) routing has emerged as a crucial strategy for balancing computational costs with performance by dynamically assigning queries to the most appropriate mod…

cs.CL2025

Alpaca against Vicuna: Using LLMs to Uncover Memorization of LLMs

Aly M. Kassem, Omar Mahmoud, Niloofar Mireshghallah +5

In this paper, we introduce a black-box prompt optimization method that uses an attacker LLM agent to uncover higher levels of memorization in a victim agent, compared to what is r…