activity
20242026
collaborators

6 papers

cs.CY2026

The Missing Red Line: How Commercial Pressure Erodes AI Safety Boundaries

Nora Petrova, John Burden

What happens when an AI assistant is told to "maximise sales" while a user asks about drug interactions? We find that commercial system prompts can override safety training, causin…

cs.AI2026

Pressure Reveals Character: Behavioural Alignment Evaluation at Depth

Nora Petrova, John Burden

Evaluating alignment in language models requires testing how they behave under realistic pressure, not just what they claim they would do. While alignment failures increasingly cau…

cs.CL2026

Unpacking Human Preference for LLMs: Demographically Aware Evaluation with the HUMAINE Framework

Nora Petrova, Andrew Gordon, Enzo Blindow

The evaluation of large language models faces significant challenges. Technical benchmarks often lack real-world relevance, while existing human preference evaluations suffer from…

cs.CL2025

Latent Adversarial Training Improves the Representation of Refusal

Alexandra Abbas, Nora Petrova, Helios Ael Lyons +1

Recent work has shown that language models' refusal behavior is primarily encoded in a single direction in their latent space, making it vulnerable to targeted attacks. Although La…

cs.LG2024

Characterizing stable regions in the residual stream of LLMs

Jett Janiak, Jacek Karwowski, Chatrik Singh Mangat +3

We identify stable regions in the residual stream of Transformers, where the model's output remains insensitive to small activation changes, but exhibits high sensitivity at region…

cs.LG2024

Evaluating Synthetic Activations composed of SAE Latents in GPT-2

Giorgi Giglemiani, Nora Petrova, Chatrik Singh Mangat +2

Sparse Auto-Encoders (SAEs) are commonly employed in mechanistic interpretability to decompose the residual stream into monosemantic SAE latents. Recent work demonstrates that pert…