collaborators

13 papers

cs.CL2026

BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

Jophin John, Michael Hoffmann, Jan Fillies +2

Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introdu…

cs.AI2026

AIMO Interpretability Challenge

Michal Štefánik, Philipp Mondorf, Andreas Waldis +11

The paper introduces the AIMO Interpretability Challenge, a competition that evaluates whether advanced mathematical language models solve olympiad‑level problems using robust reas…

cs.LG2026

ExPLAIND: Unifying Model, Data, and Training Attribution to Study Model Behavior

Florian Eichin, Yupei Du, Philipp Mondorf +3

Post-hoc interpretability methods typically attribute a model's behavior to its components, data, or training trajectory in isolation, and are often tied to a particular level of g…

cs.CL2026

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

Xinyuan Cheng, Beiduo Chen, Philipp Mondorf +1

Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, these traces can be passed to o…

cs.LG2026

Tracing Uncertainty in Language Model "Reasoning"

Nils Grünefeld, Bertram Højer, Philipp Mondorf +5

Language model (LM) "reasoning", commonly described as Chain-of-Thought or test-time scaling, often improves benchmark performance, but the dynamics underlying this process remain…

cs.CL2026

If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models

Jasmin Orth, Philipp Mondorf, Barbara Plank

Conditional acceptability refers to how plausible a conditional statement is perceived to be. It plays an important role in communication and reasoning, as it influences how indivi…