10 citations · 10 across the 3 of their papers we have counts for
3 papers
ContextBench: Modifying Contexts for Targeted Latent Activation
Robert Graham, Edward Stevinson, Leo Richter +3
Identifying inputs that trigger specific behaviours or latent features in language models could have a wide range of safety use cases. We investigate a class of methods capable of…
Gradient Routing: Masking Gradients to Localize Computation in Neural Networks
Alex Cloud, Jacob Goldman-Wetzler, Evžen Wybitul +2
Neural networks are trained primarily based on their inputs and outputs, without regard for their internal mechanisms. These neglected mechanisms determine properties that are crit…
Adversarial Policies Beat Superhuman Go AIs
Tony T. Wang, Adam Gleave, Tom Tseng +8
We attack the state-of-the-art Go-playing AI system KataGo by training adversarial policies against it, achieving a >97% win rate against KataGo running at superhuman settings. Our…