collaborators

6 papers

cs.CL2026

Inertia in Moral and Value Judgments of Large Language Models

Bruce W. Lee, Yeongheon Lee, Hyunsoo Cho

Large Language Models (LLMs) behave non-deterministically, and prompting has become a common method for steering their outputs. A popular strategy is to assign a persona to the mod…

cs.AI2026

Reasoning Models Struggle to Control their Chains of Thought

Chen Yueh-Han, Robert McCarthy, Bruce W. Lee +5

Chain-of-thought (CoT) monitoring is a promising tool for detecting misbehaviors and understanding the motivations of modern reasoning models. However, if models can control what t…

cs.LG2026

Training Agents to Self-Report Misbehavior

Bruce W. Lee, Chen Yueh-Han, Tomek Korbak

Frontier AI agents may pursue hidden goals while concealing their pursuit from oversight. Alignment training aims to prevent such behavior by reinforcing the correct goals, but ali…

cs.LG2025

Distillation Robustifies Unlearning

Bruce W. Lee, Addie Foote, Alex Infanger +6

Current LLM unlearning methods are not robust. A few steps of finetuning can revert their effects. We begin by showing that this is true even for an idealized form of unlearning: t…

cs.RO2025

DEMONSTRATE: Zero-shot Language to Robotic Control via Multi-task Demonstration Learning

Rahel Rickenbach, Bruce Lee, René Zurbrügg +2

The integration of large language models (LLMs) with control systems has demonstrated significant potential in various settings, such as task completion with a robotic manipulator.…

cs.LG2025

Programming Refusal with Conditional Activation Steering

Bruce W. Lee, Inkit Padhi, Karthikeyan Natesan Ramamurthy +4

LLMs have shown remarkable capabilities, but precisely controlling their response behavior remains challenging. Existing activation steering methods alter LLM behavior indiscrimina…