5 papers
Out-of-Distribution Detection Methods Answer the Wrong Questions
Yucen Lily Li, Daohan Lu, Polina Kirichenko +4
To detect distribution shifts and improve model safety, many out-of-distribution (OOD) detection methods rely on the predictive uncertainty or features of supervised models trained…
Utility Engineering: Analyzing and Controlling Emergent Value Systems in AIs
Mantas Mazeika, Xuwang Yin, Rishub Tamirisa +8
As AIs rapidly advance and become more agentic, the risk they pose is governed not only by their capabilities but increasingly by their propensities, including goals and values. Tr…
A Closer Look at System Prompt Robustness
Norman Mu, Jonathan Lu, Michael Lavery +1
System prompts have emerged as a critical control surface for specifying the behavior of LLMs in chat and agent settings. Developers depend on system prompts to specify important c…
Generative AI Security: Challenges and Countermeasures
Banghua Zhu, Norman Mu, Jiantao Jiao +1
Generative AI's expanding footprint across numerous industries has led to both excitement and increased scrutiny. This paper delves into the unique security challenges posed by Gen…
Mark My Words: Analyzing and Evaluating Language Model Watermarks
Julien Piet, Chawin Sitawarin, Vivian Fang +2
The capabilities of large language models have grown significantly in recent years and so too have concerns about their misuse. It is important to be able to distinguish machine-ge…