14 papers
AIMO Interpretability Challenge
Michal Štefánik, Philipp Mondorf, Andreas Waldis +11
The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…
ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior
Florian Eichin, Yupei Du, Philipp Mondorf +3
Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of g…
Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models
Xinyuan Cheng, Beiduo Chen, Philipp Mondorf +1
Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, these traces can be passed to o…
LPDS: Evaluating LLM Robustness Through Logic-Preserving Difficulty Scaling
Philipp Mondorf, Samuel J. Bell, Jesse Dodge +1
As large language models (LLMs) are increasingly deployed to perform tasks with minimal human oversight, it is crucial that these models operate robustly. In particular, a model th…
Tracing Uncertainty in Language Model "Reasoning"
Nils Grünefeld, Bertram Højer, Philipp Mondorf +5
Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain…
Language Models Learn Universal Representations of Numbers and Here's Why You Should Care
Michal Å tefánik, Timothee Mickus, Marek KadlÄÃk +7
Prior work has shown that large language models (LLMs) often converge to accurate input embedding for numbers, based on sinusoidal representations. In this work, we quantify that t…