2 citations · 2 across the 5 of their papers we have counts for
Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
REDAgentBench: Executable Red Teaming and Faithful Measurement of LLM Agent Systems
Zixing Chen, Xingyuan Liu, Jie Zhu +6
Large language model (LLM) agents combine language-based reasoning with external tools to perform complex tasks. Adversarial inputs can exploit interactions between the agent and i…
cs.AI2026
FinMCP-Bench: Benchmarking LLM Agents for Real-World Financial Tool Use under the Model Context Protocol
Jie Zhu, Yimin Tian, Boyang Li +8
This paper introduces \textbf{FinMCP-Bench}, a novel benchmark for evaluating large language models (LLMs) in solving real-world financial problems through tool invocation of finan…
cs.AI2025
DianJin-R1: Evaluating and Enhancing Financial Reasoning in Large Language Models
Jie Zhu, Qian Chen, Huaixia Dou +4
Effective reasoning remains a core challenge for large language models (LLMs) in the financial domain, where tasks often require domain-specific knowledge, precise numerical calcul…