collaborators

26 papers

cs.CC2026

How to Verify Consistency of Probabilistic Claims

Orr Paradise, Oliver Richardson, Yoshua Bengio +1

When a probabilistic predictor answers many conditional-probability queries, are its answers self-consistent, and can this be verified in polynomial time? This problem is of intere…

cs.AI2026

Safety from Honesty in a Disinterested AI Predictor

Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…

cs.CY2026

The 2026 Singapore Consensus on Global AI Safety Research Priorities

Stephen Casper, Oskar Galeev, Yoshua Bengio +117

Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…

cs.CY2026

Open Technical Problems in Open-Weight AI Model Risk Management

Stephen Casper, Kyle O'Brien, Shayne Longpre +19

Frontier AI models with openly available weights are steadily becoming more powerful and widely adopted. However, compared to proprietary models, open-weight models pose different…

cs.AI2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov +4

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untru…

cs.CL2026

When Benign Inputs Lead to Severe Harms: Eliciting Unsafe Unintended Behaviors of Computer-Use Agents

Jaylen Jones, Zhehao Zhang, Yuting Ning +6

Although computer-use agents (CUAs) hold significant potential to automate increasingly complex OS workflows, they can demonstrate unsafe unintended behaviors that deviate from exp…