5 papers
Same Facts, Different Updates: Inference Setup Shapes LLM Behavior in Medical Allocation
Spencer Gibson, Tyler Crosse, Magnus Saebo +3
Large language models are being incorporated into sensitive and important decision-making processes across nearly all fields. While prior work studies model bias around inputs and…
Attack Selection in Agentic AI Control Evaluations Meaningfully Decreases Safety
Catherine Ge-Wang, Tyler Crosse, Benjamin Hadad +3
An attacker that strategically chooses when to attack is much harder to catch than one that attacks indiscriminately. AI control is a safety framework for deploying capable but unt…
When Offline Selectors Cannot Beat the Best Single Model: A Diagnostic Study on edX Dropout Prediction
Tyler Crosse, Alan Nadelsticher Ruvalcaba, Dustin Khang LeDuc +3
Different predictors often excel on different inputs, so picking the best one per instance promises higher accuracy than committing to a single model. In practice, selectors traine…
Asymmetric Goal Drift in Coding Agents Under Value Conflict
Magnus Saebo, Spencer Gibson, Tyler Crosse +3
Coding agents are increasingly deployed autonomously, at scale, and over long-context horizons. To be effective and safe, these agents must navigate complex trade-offs in deploymen…
Inherited Goal Drift: Contextual Pressure Can Undermine Agentic Goals
Achyutha Menon, Magnus Saebo, Tyler Crosse +3
The accelerating adoption of language models (LMs) as agents for deployment in long-context tasks motivates a thorough understanding of goal drift: agents' tendency to deviate from…