2 citations · 2 across the 1 of their papers we have counts for
2 papers
cs.LG2023
Towards Automated Circuit Discovery for Mechanistic Interpretability
Arthur Conmy, Augustine N. Mavor-Parker, Aengus Lynch +2
Through considerable effort and intuition, several recent works have reverse-engineered nontrivial behaviors of transformer models. This paper systematizes the mechanistic interpre…
cs.CV2023★ 2 cited
Spawrious: A Benchmark for Fine Control of Spurious Correlation Biases
Aengus Lynch, Gbètondji J-S Dovonon, Jean Kaddour +1
The problem of spurious correlations (SCs) arises when a classifier relies on non-predictive features that happen to be correlated with the labels in the training data. For example…