From the 1 of 10 linked papers with an AI index.
4 papers · 1 filter
CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation
Mengting Chen, Yanshu Sun, Wanting Liang +5
Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge…
FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality
Beidi Luan, Rui Sun, Sinuo Wang +5
The paper introduces a scalable pipeline that automatically creates and evaluates rubrics for assessing the quality of long-form financial reports generated by deep research agents…
AdaTIR: Adaptive Tool-Integrated Reasoning via Difficulty-Aware Policy Optimization
Zhaiyu Fang, Ruipeng Sun
Tool-Integrated Reasoning (TIR) has significantly enhanced the capabilities of Large Language Models (LLMs), yet current agents tend to exhibit cognitive offloading, redundantly in…
FinResearchBench: A Logic Tree based Agent-as-a-Judge Evaluation Framework for Financial Research Agents
Rui Sun, Zuo Bai, Wentao Zhang +4
Recently, AI agents are rapidly evolving in intelligence and widely used in professional research applications, such as STEM, software development, and finance. Among these AI agen…