Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Improving Evaluation Realism with Inference-Time Compute and Deployment Scaffolds
Axel Ahlqvist, Richard Guan, Juan-Pablo Rivera +6
A core obstacle to alignment evaluation is evaluation awareness: capable models can tell when they are being tested rather than deployed, weakening the conclusions a safety evaluat…
cs.AI2026
Evaluating whether AI models would sabotage AI safety research
Robert Kirk, Alexandra Souly, Kai Fronsdal +2
We evaluate the propensity of frontier models to sabotage or refuse to assist with safety research when deployed as AI research agents within a frontier AI company. We apply two co…
cs.AI2024
MISR: Measuring Instrumental Self-Reasoning in Frontier Models
Kai Fronsdal, David Lindner
We propose a suite of tasks to evaluate the instrumental self-reasoning ability of large language model (LLM) agents. Instrumental self-reasoning ability could improve adaptability…