collaborators
Showing cs.CLShow all

6 papers · 1 filter

cs.CL2026

From Repetition to Recognition: Inductive Discovery of Disinformation Narratives

Max Upravitelev, Veronika Solopova, Jing Yang +4

In disinformation datasets, narratives are often understood as recurring interpretive patterns that group texts under narrative labels. Recent work formalized narrative mining as i…

cs.CL2026

Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

Antonin Poché, Fanny Jourdan, Nils Feldhus +6

Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is co…

cs.CL2026

Judge Circuits Explain Format-Induced Inconsistency in LLM-as-a-Judge

Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10

LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…

cs.CL2026

Macro: Enhancing Multilingual Counterfactual Explanations through Alignment-as-Preference Optimization

Yilong Wang, Qianli Wang, Bohao Chu +3

Self-generated counterfactual explanations (SCEs) are minimally modified inputs (minimality) generated by large language models (LLMs) that flip their own predictions (validity), o…

cs.CL2026

Multiperspectivity as a Resource for Narrative Similarity Prediction

Max Upravitelev, Veronika Solopova, Jing Yang +4

Predicting narrative similarity can be understood as an inherently interpretive task: different, equally valid readings of the same text can produce divergent interpretations and t…

cs.CL2026

What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025

Jing Yang, Nils Feldhus, Salar Mohtaj +10

As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) adoption has accelerated rapidly,…