Showing 2026Show all
2 papers · 1 filter
cs.AI2026
Business Arena: Benchmarking LLM Agents in a Realistic Marketplace
Yijun Pan, Yukun Lian, Kunyu Shi +5
Running a business is a challenging form of intelligent work. Operators must infer opportunities from partial signals, commit capital under uncertainty, adapt to delayed outcomes i…
cs.AI2026
JADE: Expert-Grounded Dynamic Evaluation for Open-Ended Professional Tasks
Lanbo Lin, Jiayao Liu, Tianyuan Yang +5
Evaluating agentic AI on open-ended professional tasks faces a fundamental dilemma between rigor and flexibility. Static rubrics provide rigorous, reproducible assessment but fail…