3 papers
cs.CR2026
FinVault: Benchmarking Financial Agent Safety in Execution-Grounded Environments
Zhi Yang, Runguo Li, Qiqi Qiang +15
Financial agents powered by large language models (LLMs) are increasingly deployed for investment analysis, risk assessment, and automated decision-making, where their abilities to…
cs.CE2025
VisFinEval: A Scenario-Driven Chinese Multimodal Benchmark for Holistic Financial Understanding
Zhaowei Liu, Xin Guo, Haotian Xia +11
Multimodal large language models (MLLMs) hold great promise for automating complex financial analysis. To comprehensively evaluate their capabilities, we introduce VisFinEval, the…
cs.CL2025
FinGAIA: A Chinese Benchmark for AI Agents in Real-World Financial Domain
Lingfeng Zeng, Fangqi Lou, Zixuan Wang +18
The booming development of AI agents presents unprecedented opportunities for automating complex tasks across various domains. However, their multi-step, multi-tool collaboration c…