1 paper
Andrzej Szablewski, Gabriel Konar-Steenberg, Raffaello Fornasiere +2
Model organisms (MOs) - language models trained to exhibit undesired or unnatural behaviours - are frequently used as testbeds for evaluating white-box interpretability techniques.…