From the 1 of 3 linked papers with an AI index.
3 papers
AIMO Interpretability Challenge
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…
Attend or Perish: Benchmarking Attention in Algorithmic Reasoning
Michal Spiegel, Michal Å tefánik, Marek KadlÄÃk +1
Can transformers learn to perform algorithmic tasks reliably across previously unseen input/output domains? While pre-trained language models show solid accuracy on benchmarks inco…
Negation: A Pink Elephant in the Large Language Models' Room?
Tereza Vrabcová, Marek KadlÄÃk, Petr Sojka +2
Negations are key to determining sentence meaning, making them essential for logical reasoning. Despite their importance, negations pose a substantial challenge for large language…