8 papers
Efficient Safety Alignment of Language Models via Latent Personality Traits
Mohamed Amine Merzouk, Nolan Smyth, Damiano Fornasiere +3
Current safety methods for large language models are known to be vulnerable to adversarial attacks, motivating research into robust alternatives. Latent Adversarial Training (LAT)…
Do LLMs Hold Their Values? MANTA: A Multi-Turn Adversarial Benchmark for Animal Welfare Reasoning
Isabella Luong, Joyee Chen, Arturs Kanepajs +5
Evaluating animal welfare reasoning in LLMs remains an open challenge despite rapid deployment in consumer and professional contexts where welfare considerations appear implicitly…
"I've Seen How This Goes": Characterizing Diversity via Progressive Conditional Surprise
Matthew Khoriaty, David Williams-King, Shi Feng
Measuring the diversity of creative outputs is central to evaluating post-training mode collapse, comparing decoding strategies, and quantifying creative behavior in both AI and hu…
Behavioural Analysis of Alignment Faking
Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1
Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understandi…
FragBench: Cross-Session Attacks Hidden in Benign-Looking Fragments
Astha Mehta, Niruthiha Selvanayagam, Cedric Lam +10
An attacker can split a malicious goal into sub-prompts that each look benign on their own and only become harmful in combination. Existing LLM safety benchmarks evaluate prompts o…
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
Linh Le, David Williams-King, Mohamed Amine Merzouk +2
Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerab…