activity
20192026
most citedEnsemble Model Patching: A Parameter-Efficient Variational Bayesian Neural Network

5 citations · 10 across the 16 of their papers we have counts for

collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

GPS-Bench: A Governance Policy Benchmark for Automating Policy Analysis

Linh Le, Melanie Bui, My Chiffon Nguyen +2

Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows. LLM-based poli…

cs.AI2026

Latent Personality Alignment: Improving Harmlessness Without Mentioning Harms

Linh Le, David Williams-King, Mohamed Amine Merzouk +2

Current adversarial robustness methods for large language models require extensive datasets of harmful prompts (thousands to hundreds of thousands of examples), yet remain vulnerab…

cs.AI2026

Behavioural Analysis of Alignment Faking

Nathaniel Mitrani Hadida, Rhea Karty, David Williams-King +1

Alignment faking (AF) refers to a model strategically complying with a training objective to avoid behavioural modification while preserving its deployment preferences. Understandi…

cs.AI2026

Diagnosing Pathological Chain-of-Thought in Reasoning Models

Manqing Liu, David Williams-King, Ida Caspary +5

Chain-of-thought (CoT) reasoning is fundamental to modern LLM architectures and represents a critical intervention point for AI safety. However, CoT reasoning may exhibit failure m…

cs.AI2025

Representation Engineering for Large-Language Models: Survey and Research Challenges

Lukasz Bartoszcze, Sarthak Munshi, Bryan Sukidi +6

Large-language models are capable of completing a variety of tasks, but remain unpredictable and intractable. Representation engineering seeks to resolve this problem through a new…

cs.AI20251 cited

Superintelligent Agents Pose Catastrophic Risks: Can Scientist AI Offer a Safer Path?

Yoshua Bengio, Michael Cohen, Damiano Fornasiere +10

The leading AI companies are increasingly focused on building generalist AI agents -- systems that can autonomously plan, act, and pursue goals across almost all tasks that humans…