works on

From the 1 of 10 linked papers with an AI index.

collaborators

10 papers

cs.CL2026

CalibratedRubric: Task-Adaptive Rubric Banks for Open-Ended LLM Evaluation

Mengting Chen, Yanshu Sun, Wanting Liang +5

Reliable evaluation of open-ended LLM outputs requires fine-grained rubrics, yet expert curation is costly and difficult to scale. Existing automated pipelines rely on strict judge…

cs.LG2026

GAUGE: Grading Agent-Built Financial Models Without a Golden Answer

Jiacheng Lu, Sinuo Wang, Wentao Zhao +12

Financial models combine public disclosures with analyst assumptions to produce forecasts and valuations. While some components can be checked mechanically, forecasts, discount rat…

cs.CL2026

FinResearchBench II: A Deep Research Benchmark with Consensus-Derived Gold Rubrics for Distinguishing Financial Report Quality

Beidi Luan, Rui Sun, Sinuo Wang +5

The paper introduces a scalable pipeline that automatically creates and evaluates rubrics for assessing the quality of long-form financial reports generated by deep research agents…

q-fin.TR2026

ContestTrade: A Multi-Agent Trading System Based on Internal Contest Mechanism

Rui Sun, Li Zhao, Zuoyou Jiang +5

In financial trading, large language model (LLM)-based agents demonstrate significant potential, but their decisions can be sensitive to noisy and non-stationary market information…

cs.AI2026

From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets

Taojie Zhu, Wentao Zhao, Rui Sun +7

Evaluating whether large language model (LLM) agents can profit in capital markets is increasingly framed as end-to-end trading: place an agent in a historical market, let it trade…

econ.TH2026

State-Robust Nash Predictions In Population Games

Rui Sun, Junfei Guo

This paper introduces state-robust equilibrium (SRE), a local validity test for Nash predictions in finite-strategy population games when the payoff-relevant aggregate state may be…