4 papers
Prefill Awareness in Large Language Models
Andy Wang, Parv Mahajan, David Demitri Africa +3
Safety-relevant studies of language models, including alignment and jailbreaking evaluations and AI control protocols, often rely on prefilling model outputs. If AI models can reco…
How Professional Visual Artists are Negotiating Generative AI in the Workplace
Harry H. Jiang, Jordan Taylor, William Agnew
Generative AI has been heavily critiqued by artists in both popular media and HCI scholarship. However, more work is needed to understand the impacts of generative AI on profession…
AI Failure Loops in Devalued Work: The Confluence of Overconfidence in AI and Underconfidence in Worker Expertise
Anna Kawakami, Jordan Taylor, Sarah Fox +2
A growing body of literature has focused on understanding and addressing workplace AI design failures. However, past work has largely overlooked the role of the devaluation of work…
Auditing Games for Sandbagging
Jordan Taylor, Sid Black, Dillon Bowen +10
Future AI systems could conceal their capabilities ('sandbagging') during evaluations, potentially misleading developers and auditors. We stress-tested sandbagging detection techni…