4 papers
Hidden in Thought: Transferable Chain-of-Thought Artifacts Induce Harmful Behavior
Ali khalil, Aly M. Kassem, Mohamed Abdelrazek +3
We investigate whether harmful chain-of-thought (CoT) traces from compromised language models can transfer unsafe behaviour and be distilled into reusable jailbreak attacks. Using…
Shared Latent Structures Enable Unified Backdoor Detection and Mitigation in LLMs
Omar Mahmoud, Aly M. Kassem, Thommen George Karimpanal +4
Backdoor attacks in large language models (LLMs) are often treated as isolated trigger-response failures, motivating defenses tailored to specific triggers or behaviors. We show th…
The Unintended Trade-off of AI Alignment:Balancing Hallucination Mitigation and Safety in LLMs
Omar Mahmoud, Ali Khalil, Buddhika Laknath Semage +2
Hallucination in large language models (LLMs) has been widely studied in recent years, with progress in both detection and mitigation aimed at improving truthfulness. Yet, a critic…
Improving Multilingual Language Models by Aligning Representations through Steering
Omar Mahmoud, Buddhika Laknath Semage, Thommen George Karimpanal +1
This paper investigates how Large Language Models (LLMs) represent non-English tokens -- a question that remains underexplored despite recent progress. We propose a lightweight int…