works on

From the 1 of 22 linked papers with an AI index.

activity
20242026
collaborators

21 papers

cs.CR2026

AI Security Leaderboard: Methodology, Results and Minimal Standard

Jasper Timm, Lukas Struppek, Ziwei Xu +12

The AI Security Leaderboard is an independent benchmark that ranks the safeguards of frontier AI models from least to most secure. It tests models against the FARAI Minimal Stan…

cs.AI2026

Large language models can effectively convince people to believe conspiracies

Thomas H. Costello, Kellin Pelrine, Matthew Kowal +6

The paper investigates whether large language models can be used to persuade people to adopt or reject conspiracy beliefs, finding that LLMs can both increase and decrease belief d…

cs.CY2026

The 2026 Singapore Consensus on Global AI Safety Research Priorities

Stephen Casper, Oskar Galeev, Yoshua Bengio +117

Frontier AI capabilities and autonomy are advancing rapidly. A growing number of real-world incidents make a trusted AI ecosystem essential to embracing AI with confidence. The 202…

cs.AI2026

Scaling Trends for Lie Detector Oversight in Preference Learning

Oskar J. Hollinsworth, Ann-Kathrin Dombrowski, Sam Adam-Day +2

Deceptive behavior in LLMs is costly to monitor and prevent, motivating approaches such as Scalable Oversight via Lie Detectors (SOLiD) (Cundy & Gleave, 2025), which uses lie detec…

cs.CL2026

Evaluating using Mock Tool Calls to Quarantine Untrusted Prompt Inputs

David Gros, Adam Gleave

Large language models must frequently process untrusted inputs, such as judging an answer from another model or running tasks like spam and harm classifiers while under adversarial…

cs.LG2026

The Obfuscation Atlas: Mapping Where Honesty Emerges in RLVR with Deception Probes

Mohammad Taufeeque, Stefan Heimersheim, Adam Gleave +1

Training against white-box deception detectors has been proposed as a way to make AI systems honest. However, such training risks models learning to obfuscate their deception to ev…