2 papers
cs.CL2025
MastermindEval: A Simple But Scalable Reasoning Benchmark
Jonas Golde, Patrick Haller, Fabio Barth +1
Recent advancements in large language models (LLMs) have led to remarkable performance across a wide range of language understanding and mathematical tasks. As a result, increasing…
cs.CL2025
Familiarity: Better Evaluation of Zero-Shot Named Entity Recognition by Quantifying Label Shifts in Synthetic Training Data
Jonas Golde, Patrick Haller, Max Ploner +3
Zero-shot named entity recognition (NER) is the task of detecting named entities of specific types (such as 'Person' or 'Medicine') without any training examples. Current research…