From the 1 of 22 linked papers with an AI index.
21 papers
AI Security Leaderboard: Methodology, Results and Minimal Standard
Jasper Timm, Lukas Struppek, Ziwei Xu +12
The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FARAI Minimal Stan…
Large language models can effectively convince people to believe conspiracies
Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6
The paper investigates whether large language models can be used to persuade people to adopt or reject conspiracy beliefs, finding that LLMs can both increase and decrease belief d…
The 2026 Singapore Consensus on Global AI Safety Research Priorities
Stephen Casper, Oskar Galeev, Yoshua Bengio +117
Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…
Scaling Trends for Lie Detector Oversight in Preference Learning
Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day +2
Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detec…
Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs
David Gros, Adam Gleave
Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classifiers while under adversarial…
The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes
Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave +1
Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to ev…