automated evaluation 1benchmarking 1financial report generation 1large language models 1rubric evaluation 1
From the 1 of 4 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2026
FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
Beidi Luan, Rui Sun, Sinuo Wang +5
The paper introduces a scalable pipeline that automatically creates and evaluates rubrics for assessing the quality of long-form financial reports generated by deep research agents…
cs.CL2026
O-Researcher: An Open Ended Deep Research Model via Multi-Agent Distillation and Agentic RL
Yi Yao, He Zhu, Piaohong Wang +12
The performance gap between closed-source and open-source large language models (LLMs) is largely attributed to disparities in access to high-quality training data. To bridge this…