1 paper
Daniel Kerrigan, Brian Barr, Enrico Bertini
As language models (LMs) rise in prominence, there is interest in making them more transparent in order to better understand their internal behavior. Recent interpretability work h…