Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
Mengting Chen, Yanshu Sun, Wanting Liang +5
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge…
cs.CL2026
FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
Beidi Luan, Rui Sun, Sinuo Wang +5
Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human experts to define and execu…