1 paper
Sanjay Basu, Sadiq Y. Patel, Parth Sheth +5
Language models encode task-relevant knowledge in internal representations that far exceeds their output performance, but whether mechanistic interpretability methods can bridge th…