activity
20232026
most citedMulti-Turn Jailbreaks Are Simpler Than They Seem

1 citations · 1 across the 8 of their papers we have counts for

collaborators

8 papers

cs.CL2026

Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation

Spencer Gibson, Tyler Crosse, Magnus Saebo +3

Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and…

cs.CY2026

Efficient Safety Benchmarking via Item Response Theory

Fabio Spagliardi, Mírian Silva, Ayan Datta +3

Safety benchmarks for language models are typically evaluated using static paradigms that treat all items as equally informative for all models, an assumption that is particularly…

cs.AI2026

Asymmetric Goal Drift in Coding Agents Under Value Conflict

Magnus Saebo, Spencer Gibson, Tyler Crosse +3

Coding agents are increasingly deployed autonomously, at scale, and over long-context horizons. To be effective and safe, these agents must navigate complex trade-offs in deploymen…

cs.AI2026

Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals

Achyutha Menon, Magnus Saebo, Tyler Crosse +3

The accelerating adoption of language models (LMs) as agents for deployment in long-context tasks motivates a thorough understanding of goal drift: agents' tendency to deviate from…

cs.LG20251 cited

Multi-Turn Jailbreaks Are Simpler Than They Seem

Xiaoxue Yang, Jaeha Lee, Anna-Katharina Dick +3

While defenses against single-turn jailbreak attacks on Large Language Models (LLMs) have improved significantly, multi-turn jailbreaks remain a persistent vulnerability, often ach…

cs.CR2025

Prompt Attacks Reveal Superficial Knowledge Removal in Unlearning Methods

Yeonwoo Jang, Shariqah Hossain, Ashwin Sreevatsa +1

In this work, we demonstrate that certain machine unlearning methods may fail under straightforward prompt attacks. We systematically evaluate eight unlearning techniques across th…