4 papers
Who Wins Where? Conformal Model Comparison for Local Superiority
Yi Zhou, Baishi Li, Xuan Yao +1
Standard model comparison is global, aggregating losses across the covariate space to declare a single winner. This can obscure heterogeneous performance, where different models ar…
FinDeepForecast: A Live Multi-Agent System for Benchmarking Deep Research Agents in Financial Forecasting
Xiangyu Li, Xuan Yao, Guohao Qi +16
Deep Research (DR) Agents powered by advanced Large Language Models (LLMs) have fundamentally shifted the paradigm for completing complex research tasks. Yet, a comprehensive and l…
FinDeepResearch: Evaluating Deep Research Agents in Rigorous Financial Analysis
Fengbin Zhu, Xiang Yao Ng, Ziyang Liu +19
Deep Research (DR) agents, powered by advanced Large Language Models (LLMs), have recently garnered increasing attention for their capability in conducting complex research tasks.…
Evaluating Large Language Models for Financial Reasoning: A CFA-Based Benchmark Study
Xuan Yao, Qianteng Wang, Xinbo Liu +1
The rapid advancement of large language models presents significant opportunities for financial applications, yet systematic evaluation in specialized financial contexts remains li…