2 papers
cs.LG2026
Fuzzing Large Language Models to Elicit Hidden Behaviours
Mohammed Abu Baker, Lakshmi Babu-Saheer
Sleeper agents are the canonical model organism of deception: models trained to behave normally but to emit an unsafe behaviour on a specific trigger. Eliciting that behaviour with…
cs.CL2025
Mechanistic Exploration of Backdoored Large Language Model Attention Patterns
Mohammed Abu Baker, Lakshmi Babu-Saheer
Backdoor attacks creating 'sleeper agents' in large language models (LLMs) pose significant safety risks. This study employs mechanistic interpretability to explore resulting inter…