9 citations · 9 across the 1 of their papers we have counts for
3 papers
Structured World Representations in Maze-Solving Transformers
Michael Igorevich Ivanitskiy, Alex F. Spies, Tilman Räuker +9
Transformer models underpin many recent advances in practical machine learning applications, yet understanding their internal behavior continues to elude researchers. Given the siz…
Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
Rusheb Shah, Quentin Feuillade--Montixi, Soroush Pour +3
Despite efforts to align large language models to produce harmless responses, they are still vulnerable to jailbreak prompts that elicit unrestricted behaviour. In this work, we in…
A Configurable Library for Generating and Manipulating Maze Datasets
Michael Igorevich Ivanitskiy, Rusheb Shah, Alex F. Spies +8
Understanding how machine learning models respond to distributional shifts is a key research challenge. Mazes serve as an excellent testbed due to varied generation algorithms offe…