23 citations · 26 across the 4 of their papers we have counts for
4 papers
The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
Christina Lu, Jack Gallagher, Jonathan Michala +2
Large language models can represent a variety of personas but typically default to a helpful Assistant identity cultivated during post-training. We investigate the structure of the…
The LLM Has Left The Chat: Evidence of Bail Preferences in Large Language Models
Danielle Ensign, Henry Sleight, Kyle Fish
When given the option, will LLMs choose to leave the conversation (bail)? We investigate this question by giving models the option to bail out of interactions using three different…
Will AI Tell Lies to Save Sick Children? Litmus-Testing AI Values Prioritization with AIRiskDilemmas
Yu Ying Chiu, Zhilin Wang, Sharan Maiya +4
Detecting AI risks becomes more challenging as stronger models emerge and find novel methods such as Alignment Faking to circumvent these detection attempts. Inspired by how risky…
Taking AI Welfare Seriously
Robert Long, Jeff Sebo, Patrick Butlin +7
In this report, we argue that there is a realistic possibility that some AI systems will be conscious and/or robustly agentic in the near future. That means that the prospect of AI…