10 papers
Who Annotates in NLP? A Large-scale Assessment of Human Annotation Reporting between 2018 and 2025
Maria Kunilovskaya, Gagan Bhatia, Lisa Sophie Albertelli +10
Human annotation is the empirical foundation of much NLP research, from dataset construction to model evaluation, but papers often leave unclear who produced the annotations and ho…
ValueGround: Evaluating Culture-Conditioned Visual Value Grounding in MLLMs
Zhipin Wang, Christoph Leiter, Christian Frey +3
Cultural values are expressed not only through language but also through visual scenes and everyday social practices. Yet existing evaluations of cultural values in language models…
ExCAM: Explainable Cultural Awareness Metrics
Christoph Leiter, Haiyue Song, Hour Kaing +4
Evaluating the cultural awareness of large language models is crucial to ensure the fairness of generated text and the generalizability of applications across the world. Recent ben…
GerAV: Towards New Heights in German Authorship Verification using Fine-Tuned LLMs on a New Benchmark
Lotta Kiefer, Christoph Leiter, Sotaro Takeshita +2
Authorship verification (AV) is the task of determining whether two texts were written by the same author and has been studied extensively, predominantly for English data. In contr…
CROC: Evaluating and Training T2I Metrics with Pseudo- and Human-Labeled Contrastive Robustness Checks
Christoph Leiter, Yuki M. Asano, Margret Keuper +1
The assessment of evaluation metrics (meta-evaluation) is crucial for determining the suitability of existing metrics in text-to-image (T2I) generation tasks. Human-based meta-eval…
DeepSeek-R1 vs. o3-mini: How Well can Reasoning LLMs Evaluate MT and Summarization?
Daniil Larionov, Sotaro Takeshita, Ran Zhang +5
Reasoning-enabled large language models (LLMs) excel in logical tasks, yet their utility for evaluating natural language generation remains unexplored. This study systematically co…