18 papers
What Are We Measuring in NLG? A Meta-Analysis of Evaluation Trends 2020-2025
Jing Yang, Nils Feldhus, Salar Mohtaj +10
As Natural Language Generation (NLG) dominates modern NLP, scalable evaluation remains a critical bottleneck. Consequently, LLM-as-a-judge (LaaJ) adoption has accelerated rapidly,…
Judge Circuits
Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10
LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…
Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall
Qianli Wang, Mingyang Wang, Nils Feldhus +5
Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's effects on various LLM capabilit…
Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation
Qianli Wang, Van Bach Nguyen, Yihong Liu +6
Counterfactuals refer to minimally edited inputs that cause a model's prediction to change, serving as a promising approach to explaining the model's behavior. Large language model…
Simplifying Outcomes of Language Model Component Analyses with ELIA
Aaron Louis Eidt, Nils Feldhus
While mechanistic interpretability has developed powerful tools to analyze the internal workings of Large Language Models (LLMs), their complexity has created an accessibility gap,…
Persona Prompting as a Lens on LLM Social Reasoning
Jing Yang, Moritz Hechtbauer, Elisabeth Khalilov +3
For socially sensitive tasks like hate speech detection, the quality of explanations from Large Language Models (LLMs) is crucial for factors like user trust and model alignment. W…