1 paper
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
We propose the AIMO Interpretability Challenge, a competition on distinguishing robust from spurious reasoning in frontier mathematical language models based on the models' interna…