4 papers
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
Mengting Chen, Yanshu Sun, Wanting Liang +5
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge…
GAUGE: Grading Agent-Built Financial Models Without a Golden Answer
Jiacheng Lu, Sinuo Wang, Wentao Zhao +12
Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rat…
FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
Beidi Luan, Rui Sun, Sinuo Wang +5
Deep research agents are increasingly used to produce long-form financial reports, yet large-scale evaluation remains bottlenecked by the need for human experts to define and execu…
Trade-R1: Bridging Verifiable Rewards to Stochastic Environments via Process-Level Reasoning Verification
Rui Sun, Yifan Sun, Sheng Xu +5
Reinforcement Learning (RL) has enabled Large Language Models (LLMs) to achieve remarkable reasoning in domains like mathematics and coding, where verifiable rewards provide clear…