works on

From the 1 of 5 linked papers with an AI index.

collaborators

5 papers

cs.AI2026

AIMO Interpretability Challenge

Michal Štefánik, Philipp Mondorf, Andreas Waldis +11

The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…

cs.CL2026

Instructions Shape Production of Language, not Processing

Andreas Waldis, Leshem Choshen, Yufang Hou +1

Instructions trigger a production-centered mechanism in language models. Through a cognitively inspired lens that separates language processing and production, we reveal this mecha…

cs.CL2026

Holmes: A Benchmark to Assess the Linguistic Competence of Language Models

Andreas Waldis, Yotam Perlitz, Leshem Choshen +2

We introduce Holmes, a new benchmark designed to assess language models (LMs) linguistic competence - their unconscious understanding of linguistic phenomena. Specifically, we use…

cs.CL2025

A Pipeline to Assess Merging Methods via Behavior and Internals

Yutaro Sigrist, Andreas Waldis

Merging methods combine the weights of multiple language models (LMs) to leverage their capacities, such as for domain adaptation. While existing studies investigate merged models…

cs.CL2025

Aligned Probing: Relating Toxic Behavior and Model Internals

Andreas Waldis, Vagrant Gautam, Anne Lauscher +2

We introduce aligned probing, a novel interpretability framework that aligns the behavior of language models (LMs), based on their outputs, and their internal representations (inte…