#automated evaluation
topicautomated evaluation
2 papers · 1 filter
cs.CL2026
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Bertil Braun, Martin Forell
The paper presents a scalable framework for automatically evaluating large language model outputs using pairwise comparisons and an Elo rating system, achieving rankings that align…
#automated evaluation#large language models#pairwise comparison#elo rating
cs.CL2026
FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
Beidi Luan, Rui Sun, Sinuo Wang +5
The paper introduces a scalable pipeline that automatically creates and evaluates rubrics for assessing the quality of long-form financial reports generated by deep research agents…
#financial report generation#benchmarking#rubric evaluation#large language models