6 citations · 10 across the 2 of their papers we have counts for
2 papers
cs.CY2024★ 6 cited
Safety Cases: How to Justify the Safety of Advanced AI Systems
Joshua Clymer, Nick Gabrieli, David Krueger +1
As AI systems become more advanced, companies and regulators will make difficult decisions about whether it is safe to train and deploy them. To prepare for these decisions, we inv…
cs.CL2023★ 4 cited
Steering Llama 2 via Contrastive Activation Addition
Nina Panickssery, Nick Gabrieli, Julian Schulz +3
We introduce Contrastive Activation Addition (CAA), an innovative method for steering language models by modifying their activations during forward passes. CAA computes "steering v…