activity
20242026
collaborators

37 papers

cs.CR2026

Stealing Reasoning Traces from Proprietary LLM APIs

Alexander Panfilov, David Schmotz, Ilia Shumailov +5

Leading large language model providers now conceal their models' step-by-step reasoning, or chain-of-thought, to protect intellectual property and limit information leakage. Rather…

cs.LG2026

Attractor States Emerge in Multi-Turn LLM Conversations

Ting-Wen Ko, Jonas Geiping

Large language models (LLMs) are increasingly used in open-ended multi-agent settings, but the long-run dynamics of model--model interaction remain poorly understood. We study whet…

stat.ML2026

When are likely answers right? On Sequence Probability and Correctness in LLMs

Johannes Zenn, Jonas Geiping

Many decoding methods for large language models can be understood as shifting probability mass toward outputs that are more likely under the model, either locally at the token leve…

cs.CL2026

Models That Know How Evaluations Are Designed Score Safer

Katharina Deckenbach, Haritz Puerto, Jonas Geiping +1

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such a…

cs.AI2026

CIAware-Bench: Benchmarking Control Intervention Awareness Across Frontier LLMs

Joachim Schaeffer, Thomas Jiralerspong, Alexander Panfilov +4

AI control protocols oversee untrusted models by monitoring their actions and modifying potentially unsafe steps, often using a trusted model. This partially tampers with the untru…

cs.LG2026

Claudini: Autoresearch Discovers State-of-the-Art Adversarial Attack Algorithms for LLMs

Alexander Panfilov, Peter Romov, Igor Shilov +3

We show that AI agents are capable of discovering novel algorithms for adversarial attacks against LLMs, advancing the state of the art on white-box jailbreaking and prompt injecti…