Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Confidence and Stability of Global and Pairwise Scores in NLP Evaluation
Georgii Levtsov, Dmitry Ustalov
With the advent of highly capable instruction-tuned neural language models, benchmarking in natural language processing (NLP) is increasingly shifting towards pairwise comparison l…
cs.CL2024
Reliable, Reproducible, and Really Fast Leaderboards with Evalica
Dmitry Ustalov
The rapid advancement of natural language processing (NLP) technologies, such as instruction-tuned large language models (LLMs), urges the development of modern evaluation protocol…