Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
BankerToolBench: Evaluating AI Agents in End-to-End Investment Banking Workflows
Elaine Lau, Markus Dücker, Ronak Chaudhary +24
Existing AI benchmarks lack the fidelity to assess economically meaningful progress on professional workflows. To evaluate frontier AI agents in a high-value, labor-intensive profe…
cs.AI2025
PAFFA: Premeditated Actions For Fast Agents
Shambhavi Krishna, Zheng Chen, Yuan Ling +4
Modern AI assistants have made significant progress in natural language understanding and tool-use, with emerging efforts to interact with Web interfaces. However, current approach…