1 citations · 1 across the 4 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Realistic honeypot evaluations for scheming propensity
Victoria Krakovna, David Lindner, Lewis Ho +2
We introduce scheming honeypot evaluations, a framework for testing whether models will pursue instrumental goals if given the opportunity. Our scheming honeypot evaluations take t…
cs.LG2025
Evaluating Frontier Models for Stealth and Situational Awareness
Mary Phuong, Roland S. Zimmermann, Ziyue Wang +6
Recent work has demonstrated the plausibility of frontier AI models scheming -- knowingly and covertly pursuing an objective misaligned with its developer's intentions. Such behavi…