2 papers
cs.AI2026
How Well Do Models Follow Their Constitutions?
Arya Jakkli, Senthooran Rajamanoharan, Neel Nanda
Frontier AI developers now train models against long written behavioral specifications, such as Anthropic's constitution (Anthropic, 2025a) and OpenAI's Model Spec (OpenAI, 2025a),…
cs.LG2026
Censored LLMs as a Natural Testbed for Secret Knowledge Elicitation
Helena Casademunt, Bartosz CywiÅski, Khoi Tran +3
Large language models sometimes produce false or misleading responses. Two approaches to this problem are honesty elicitation -- modifying prompts or weights so that the model answ…