5 citations · 10 across the 16 of their papers we have counts for
6 papers · 1 filter
GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis
Linh Le, Melanie Bui, My Chiffon Nguyen +2
Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based poli…
Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms
Linh Le, David Williams-King, Mohamed Amine Merzouk +2
Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerab…
Behavioural Analysis of Alignment Faking
Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1
Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understandi…
Diagnosing Pathological Chain-of-Thought in Reasoning Models
Manqing Liu, David Williams-King, Ida Caspary +5
Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT reasoning may exhibit failure m…
Representation Engineering for Large-Language Models: Survey and Research Challenges
Lukasz Bartoszcze, Sarthak Munshi, Bryan Sukidi +6
Large-language models are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation engineering seeks to resolve this problem through a new…
Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?
Yoshua Bengio, Michael Cohen, Damiano Fornasiere +10
The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans…