14 citations · 14 across the 5 of their papers we have counts for
1 paper · 1 filter
Jason Gross, Rajashree Agrawal, Thomas Kwa +5
We propose using mechanistic interpretability -- techniques for reverse engineering model weights into human-interpretable algorithms -- to derive and compactly prove formal guaran…