5 papers
Intertemporal Preference Steering in Qwen3 via Contrastive Activation Addition
Michal Mráz, Michal Mráz, Justin Shenk
We study linear representations of temporal horizon in the large language model Qwen3-32B and use them to change the model's time-related preferences, recommendations, and capabili…
The Knowing-Saying Gap: When Probes See Errors that Confidence Misses
Jyotin Goel, Ipshita Bandyopadhyay, Justin Shenk
Linear probes detect corrupted context in language models with near-perfect accuracy, yet this does not translate into reliable failure prediction. The result is a dissociation wit…
Temporal Preference Concepts and their Functions in a Large Language Model
Ian Rios-Sialer, Shantanu Darveshi, Shuai Jiang +4
Large Language Models (LLMs) are increasingly being deployed to make decisions that require trading off near-term gains against long-term consequences, yet little is known about ho…
Creative Collision: Directorial Persona Steering and Competition in Large Language Models
Subramanyam Sahoo, Justin Shenk
Activation steering has emerged as a powerful tool for shaping the behaviour of large language models at inference time, yet most prior work injects a \emph{single} semantic direct…
LinuxArena: A Control Setting for AI Agents in Live Production Software Environments
Tyler Tracy, Ram Potham, Nick Kuhn +31
We introduce LinuxArena, a control setting in which agents operate directly on live, multi-service production environments. LinuxArena contains 20 environments, 1,671 main tasks re…