Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Counterfactual Trace Auditing of LLM Agent Skills
Xiaolin Zhou, Jinbo Liu, Li Li +2
Large Language Model agents are increasingly augmented with agent skills. Current evaluation methods for skills remain limited. Most deployed benchmarks report only pass rate befor…
cs.AI2026
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents
Xiaolin Zhou, Aojie Yuan, Zheng Luo +12
Tool-use language agents are evaluated on benchmarks that assume clean inputs, unambiguous tool registries, and reliable APIs. Real deployments violate all these assumptions: user…