Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
MBABench: Evaluating LLM Agents on End-to-End Spreadsheet Tasks in Finance
Thomson Yen, Julian Poeltl, Harshith Srinivas Gear +10
LLM agents are increasingly expected to carry out end-to-end workflows, producing complete artifacts from high-level user instructions. To meet enterprise needs, frontier AI labs h…
cs.AI2026
SynthTools: A Framework for Scaling Synthetic Tools for Agent Development
Tommaso Castellani, Naimeng Ye, Daksh Mittal +4
For agentic systems to use external tools to solve complex, long-horizon tasks, we need a large set of diverse and controllable tool-use environments. We introduce SynthTools, a fu…