2 papers
cs.CL2026
Steering Evaluation-Aware Language Models to Act Like They Are Deployed
Tim Tian Hua, Andrew Qin, Samuel Marks +1
Large language models (LLMs) can sometimes detect when they are being evaluated and adjust their behavior to appear more aligned, compromising the reliability of safety evaluations…
cs.CY2025
Combining Cost-Constrained Runtime Monitors for AI Safety
Tim Tian Hua, James Baskerville, Henri Lemoine +3
Monitoring AIs at runtime can help us detect and stop harmful actions. In this paper, we study how to efficiently combine multiple runtime monitors into a single monitoring protoco…