Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Measuring the Behavioral Fidelity of Long-Horizon Human Activity Simulations
Yi Fei Cheng, Fan Yang, Iremsu Bas +3
As LLM-based human simulators are increasingly used for policy, evaluation, and training, they must faithfully reproduce real behavioral patterns. While prior work has examined beh…
cs.AI2025
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
Jun Takahashi, Atsunori Moteki, Akiyoshi Uchida +11
This paper introduces FieldWorkArena, a benchmark for agentic AI targeting real-world field work. With the recent increase in demand for agentic AI, they are built to detect and do…