Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Prefill Awareness in Large Language Models
Andy Wang, Parv Mahajan, David Demitri Africa +3
Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can reco…
cs.AI2025
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…