37 papers
Stealing Reasoning Traces from Proprietary LLM APIs
Alexander Panfilov, David Schmotz, Ilia Shumailov +5
Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather…
Attractor States Emerge in Multi-Turn LLM Conversations
Ting-Wen Ko, Jonas Geiping
Large language models (LLMs) are increasingly used in open-ended multi-agent settings, but the long-run dynamics of model--model interaction remain poorly understood. We study whet…
When are likely answers right? On Sequence Probability and Correctness in LLMs
Johannes Zenn, Jonas Geiping
Many decoding methods for large language models can be understood as shifting probability mass toward outputs that are more likely under the model, either locally at the token leve…
Models That Know How Evaluations Are Designed Score Safer
Katharina Deckenbach, Haritz Puerto, Jonas Geiping +1
The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such a…
CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs
Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov +4
AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untru…
Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs
Alexander Panfilov, Peter Romov, Igor Shilov +3
We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injecti…