7 papers
From Repetition to Recognition: Inductive Discovery of Disinformation Narratives
Max Upravitelev, Veronika Solopova, Jing Yang +4
In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as i…
Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
Antonin Poché, Fanny Jourdan, Nils Feldhus +6
Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is co…
Efficient Multi-objective Prompt Optimization via Pure-exploration Bandits
Donghao Li, Chengshuai Shi, Weijuan Ou +2
Prompt engineering has become central to eliciting the capabilities of large language models (LLMs). At its core lies prompt selection -- efficiently identifying the most effective…
Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge
Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10
LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…
Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization
Yilong Wang, Qianli Wang, Bohao Chu +3
Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), o…
Multiperspectivity as a Resource for Narrative Similarity Prediction
Max Upravitelev, Veronika Solopova, Jing Yang +4
Predicting narrative similarity can be understood as an inherently interpretive task: different, equally valid readings of the same text can produce divergent interpretations and t…