From the 1 of 5 linked papers with an AI index.
5 papers
AIMO Interpretability Challenge
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…
Instructions Shape Production of Language, not Processing
Andreas Waldis, Leshem Choshen, Yufang Hou +1
Instructions trigger a production-centered mechanism in language models. Through a cognitively inspired lens that separates language processing and production, we reveal this mecha…
Holmes: A Benchmark to Assess the Linguistic Competence of Language Models
Andreas Waldis, Yotam Perlitz, Leshem Choshen +2
We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic phenomena. Specifically, we use…
A Pipeline to Assess Merging Methods via Behavior and Internals
Yutaro Sigrist, Andreas Waldis
Merging methods combine the weights of multiple language models (LMs) to leverage their capacities, such as for domain adaptation. While existing studies investigate merged models…
Aligned Probing: Relating Toxic Behavior and Model Internals
Andreas Waldis, Vagrant Gautam, Anne Lauscher +2
We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (inte…