3 papers
cs.CL2026
Evaluation Pitfalls and Sparsity Limitations in LLM-based Confidence Estimates for Classification
Elena Merdjanovska, Omar Zaidan, Andreas Rücklé
Confidence estimation is essential when LLMs are used for classification, indicating when predictions can be trusted. However, common approaches such as verbalization produce extre…
cs.CL2026
Self-Aware Knowledge Probing: Evaluating Language Models' Relational Knowledge through Confidence Calibration
Christopher Kissling, Elena Merdjanovska, Alan Akbik
Knowledge probing quantifies how much relational knowledge a language model (LM) has acquired during pre-training. Existing knowledge probes evaluate model capabilities through met…
cs.CL2024
NoiseBench: Benchmarking the Impact of Real Label Noise on Named Entity Recognition
Elena Merdjanovska, Ansar Aynetdinov, Alan Akbik
Available training data for named entity recognition (NER) often contains a significant percentage of incorrect labels for entity types and entity boundaries. Such label noise pose…