activity
20232026
collaborators
Showing cs.CLShow all

18 papers · 1 filter

cs.CL2026

Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

Antonin Poché, Fanny Jourdan, Nils Feldhus +6

Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is co…

cs.CL2026

Investigating the Interplay between Contextual and Parametric Chain-of-Thought Faithfulness under Optimization

Jingyi Sun, Qianli Wang, Pepa Atanasova +2

Chain-of-Thought (CoT) faithfulness, i.e., whether CoTs genuinely reflect large language models' (LLM) underlying behavior, is typically evaluated with metrics under two disjoint p…

cs.CL2026

Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge

Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10

LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…

cs.CL2026

Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

Yilong Wang, Qianli Wang, Bohao Chu +3

Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), o…

cs.CL2026

eTracer: Towards Traceable Text Generation via Claim-Level Grounding

Bohao Chu, Qianli Wang, Hendrik Damm +5

How can system-generated responses be efficiently verified, especially in the high-stakes biomedical domain? To address this challenge, we introduce eTracer, a plug-and-play framew…

cs.CL2026

Locate, Steer, and Improve: A Practical Survey of Actionable Mechanistic Interpretability in Large Language Models

Hengyuan Zhang, Zhihao Zhang, Mingyang Wang +26

Mechanistic Interpretability (MI) has emerged as a vital approach to demystify the opaque decision-making of Large Language Models (LLMs). However, existing reviews primarily treat…