2 papers
cs.LG2026
Abstractive Red-Teaming of Language Model Character
Nate Rahn, Allison Qi, Avery Griffin +3
We want language model assistants to conform to a character specification, which asserts how the model should act across diverse user interactions. While models typically follow th…
cs.CL2024
Controlling Large Language Model Agents with Entropic Activation Steering
Nate Rahn, Pierluca D'Oro, Marc G. Bellemare
The rise of large language models (LLMs) has prompted increasing interest in their use as in-context learning agents. At the core of agentic behavior is the capacity for exploratio…