1 paper
Alex F. Spies, William Edwards, Michael I. Ivanitskiy +5
Recent studies in interpretability have explored the inner workings of transformer models trained on tasks across various domains, often discovering that these networks naturally d…