collaborators

5 papers

cs.LG2026

Mitigating Adaptive Attacks against Reasoning Models with Activation Consistency Training

Avidan Shah, Jannik Brinkmann, Rico Angell

As LLMs gain stronger reasoning capabilities, their extended chain-of-thought introduces new degrees of complexity for defending against adversarial jailbreaks and prompt injection…

cs.LG2026

On-Policy Consistency Training Improves LLM Safety with Minimal Capability Degradation

Andy Han, Kristina Fujimoto, Avidan Shah +5

Aligned models can misbehave in several ways: they are often sycophantic, fall victim to jailbreaks, or fail to include appropriate safety warnings. Consistency training is a promi…

cs.AI2026

Is It Thinking or Cheating? Detecting Implicit Reward Hacking by Measuring Reasoning Effort

Xinpeng Wang, Nitish Joshi, Barbara Plank +2

Reward hacking, where a reasoning model exploits loopholes in a reward function to achieve high rewards without solving the intended task, poses a significant threat. This behavior…

cs.LG2025

Jailbreak Transferability Emerges from Shared Representations

Rico Angell, Jannik Brinkmann, He He

Jailbreak transferability is the surprising phenomenon when an adversarial attack compromising one model also elicits harmful responses from other models. Despite widespread demons…

cs.CR2025

Monitoring Decomposition Attacks in LLMs with Lightweight Sequential Monitors

Chen Yueh-Han, Nitish Joshi, Yulin Chen +3

Current LLM safety defenses fail under decomposition attacks, where a malicious goal is decomposed into benign subtasks that circumvent refusals. The challenge lies in the existing…