1 paper · 1 filter
Chenxin Li, Zhengyang Tang, Mingxin Huang +8
LLM agents are expected to complete end-to-end units of work across software tools, business services, and local workspaces. Yet many agent benchmarks freeze a curated task set at…