4 papers
The Structural Safety Generalization Problem
Julius Broomfield, Tom Gibbs, Ethan Kosak-Hine +7
LLM jailbreaks are a widespread safety challenge. Given this problem has not yet been tractable, we suggest targeting a key failure mechanism: the failure of safety to generalize a…
Shared Autonomy for Proximal Teaching
Megha Srivastava, Reihaneh Iranmanesh, Yuchen Cui +6
Motor skill learning often requires experienced professionals who can provide personalized instruction. Unfortunately, the availability of high-quality training can be limited for…
Generating Text from Uniform Meaning Representation
Emma Markle, Reihaneh Iranmanesh, Shira Wein
Uniform Meaning Representation (UMR) is a recently developed graph-based semantic representation, which expands on Abstract Meaning Representation (AMR) in a number of ways, in par…
Emerging Vulnerabilities in Frontier Models: Multi-Turn Jailbreak Attacks
Tom Gibbs, Ethan Kosak-Hine, George Ingebretsen +6
Large language models (LLMs) are improving at an exceptional rate. However, these models are still susceptible to jailbreak attacks, which are becoming increasingly dangerous as mo…