activity
20242026
collaborators

11 papers

cs.CL2026

Judge Circuits

Nils Feldhus, Tanja Baeumel, Elena Golimblevskaia +10

LLM-as-a-judge has become the dominant paradigm for grading model outputs at scale, yet the same model assigns systematically different scores when its output format changes (e.g.,…

cs.CL2026

Through a Compressed Lens: Investigating The Impact of Quantization on Factual Knowledge Recall

Qianli Wang, Mingyang Wang, Nils Feldhus +5

Quantization methods are widely used to accelerate inference and streamline the deployment of large language models (LLMs). Although quantization's effects on various LLM capabilit…

cs.CL2026

Parallel Universes, Parallel Languages: A Comprehensive Study on LLM-based Multilingual Counterfactual Example Generation

Qianli Wang, Van Bach Nguyen, Yihong Liu +6

Counterfactuals refer to minimally edited inputs that cause a model's prediction to change, serving as a promising approach to explaining the model's behavior. Large language model…

cs.SE2026

Gendered Prompting and LLM Code Review: How Gender Cues in the Prompt Shape Code Quality and Evaluation

Lynn Janzen, Üveys Eroglu, Dorothea Kolossa +4

LLMs are increasingly embedded in programming workflows, from code generation to automated code review. Yet, how gendered communication styles interact with LLM-assisted programmin…

cs.CL2026

Can Large Language Models Still Explain Themselves? Investigating the Impact of Quantization on Self-Explanations

Qianli Wang, Nils Feldhus, Pepa Atanasova +5

Quantization is widely used to accelerate inference and streamline the deployment of large language models (LLMs), yet its effects on self-explanations (SEs) remain unexplored. SEs…

cs.CL2025

Truth or Twist? Optimal Model Selection for Reliable Label Flipping Evaluation in LLM-based Counterfactuals

Qianli Wang, Van Bach Nguyen, Nils Feldhus +4

Counterfactual examples are widely employed to enhance the performance and robustness of large language models (LLMs) through counterfactual data augmentation (CDA). However, the s…