4 citations · 6 across the 2 of their papers we have counts for
2 papers
cs.CR2024★ 2 cited
Towards evaluations-based safety cases for AI scheming
Mikita Balesni, Marius Hobbhahn, David Lindner +13
We sketch how developers of frontier AI systems could construct a structured rationale -- a 'safety case' -- that an AI system is unlikely to cause catastrophic outcomes through sc…
cs.LG2023★ 4 cited
Localizing Model Behavior with Path Patching
Nicholas Goldowsky-Dill, Chris MacLeod, Lucas Sato +1
Localizing behaviors of neural networks to a subset of the network's components or a subset of interactions between components is a natural first step towards analyzing network mec…