2 papers
cs.AI2026
ContextBench: Modifying Contexts for Targeted Latent Activation
Robert Graham, Edward Stevinson, Leo Richter +3
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…
cs.LG2025
Red-teaming Activation Probes using Prompted LLMs
Phil Blandfort, Robert Graham
Activation probes are attractive monitors for AI systems due to low cost and latency, but their real-world robustness remains underexplored. We ask: What failure modes arise under…