4 papers
APS: Bias-Controlled Adaptive Prototype Simulation for Population-Scale LLM Agents
Quan Zheng, Yan Gao, Shaobin He +6
LLM-agent simulation offers a flexible computational tool for studying population response trajectories that depend on scenario events, memory, demographics, and evolving social co…
Finch: Benchmarking Finance & Accounting across Spreadsheet-Centric Enterprise Workflows
Haoyu Dong, Pengkun Zhang, Yan Gao +6
We introduce FinWorkBench (a.k.a. Finch) for evaluating AI agents on real-world, enterprise-grade finance and accounting workflows that interleave data entry, structuring, formatti…
STRUCTUREDAGENT: Planning with AND/OR Trees for Long-Horizon Web Tasks
ELita Lobo, Xu Chen, Jingjing Meng +5
Recent advances in large language models (LLMs) have enabled agentic systems for sequential decision-making. Such agents must perceive their environment, reason across multiple tim…
GTM: Simulating the World of Tools for AI Agents
Zhenzhen Ren, Xinpeng Zhang, Zhenxing Qian +4
The integration of external tools is pivotal for empowering Large Language Model (LLM) agents with real-world capabilities. However, training these agents through direct, continuou…