activity
20242026
most citedInternational AI Safety Report 2026

1 citations · 2 across the 10 of their papers we have counts for

collaborators
Showing cs.AIShow all

7 papers · 1 filter

cs.AI2026

Safety from Honesty in a Disinterested AI Predictor

Yoshua Bengio, Oliver Richardson, Tomáš Gavenčiak +13

As AI systems become more capable, training procedures that optimize for downstream outcomes risk introducing implicit agency: goal-directed behavior that designers never specified…

cs.AI2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov +4

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untru…

cs.AI2025

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety

Tomek Korbak, Mikita Balesni, Elizabeth Barnes +38

AI systems that "think" in human language offer a unique opportunity for AI safety: we can monitor their chains of thought (CoT) for the intent to misbehave. Like all other known A…

cs.AI2025

A Definition of AGI

Dan Hendrycks, Dawn Song, Christian Szegedy +30

The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quant…

cs.AI2025

Expert-Guided LLM Reasoning for Battery Discovery: From AI-Driven Hypothesis to Synthesis and Characterization

Shengchao Liu, Hannan Xu, Yan Ai +3

Large language models (LLMs) leverage chain-of-thought (CoT) techniques to tackle complex problems, representing a transformative breakthrough in artificial intelligence (AI). Howe…

cs.AI2025

The Singapore Consensus on Global AI Safety Research Priorities

Yoshua Bengio, Tegan Maharaj, Luke Ong +84

Rapidly improving AI capabilities and autonomy hold significant promise of transformation, but are also driving vigorous debate on how to ensure that AI is safe, i.e., trustworthy,…