works on

From the 1 of 15 linked papers with an AI index.

activity
20242026
collaborators
Showing cs.CLShow all

9 papers · 1 filter

cs.CL2026

BavGround: A Benchmark for Regional Cultural Grounding and Dialect Competence in Bavarian

Jophin John, Michael Hoffmann, Jan Fillies +2

Cultural evaluation of large language models (LLMs) often focuses on high-resource standard languages, leaving regional culture and dialect communities underrepresented. We introdu…

cs.CL2026

Reasoning that Travels: Dissecting How Chain-of-Thought Transfers Across Models

Xinyuan Cheng, Beiduo Chen, Philipp Mondorf +1

Large reasoning models (LRMs) often generate extensive chain-of-thought (CoT) traces before producing a final answer. As explicit textual artifacts, these traces can be passed to o…

cs.CL2026

If Probable, Then Acceptable? Understanding Conditional Acceptability Judgments in Large Language Models

Jasmin Orth, Philipp Mondorf, Barbara Plank

Conditional acceptability refers to how plausible a conditional statement is perceived to be. It plays an important role in communication and reasoning, as it influences how indivi…

cs.CL2025

The Validation Gap: A Mechanistic Analysis of How Language Models Compute Arithmetic but Fail to Validate It

Leonardo Bertolazzi, Philipp Mondorf, Barbara Plank +1

The ability of large language models (LLMs) to validate their output and identify potential errors is crucial for ensuring robustness and reliability. However, current research ind…

cs.CL2025

MAKIEval: A Multilingual Automatic WiKidata-based Framework for Cultural Awareness Evaluation for LLMs

Raoyuan Zhao, Beiduo Chen, Barbara Plank +1

Large language models (LLMs) are used globally across many languages, but their English-centric pretraining raises concerns about cross-lingual disparities for cultural awareness,…

cs.CL2025

Probing LLMs for Multilingual Discourse Generalization Through a Unified Label Set

Florian Eichin, Yang Janet Liu, Barbara Plank +1

Discourse understanding is essential for many NLP tasks, yet most existing work remains constrained by framework-dependent discourse representations. This work investigates whether…