3 papers
cs.CL2026
ArcANE: Do Role-Playing Language Agents Stay in Character at the Right Time?
Woojung Song, Nalim Kim, Sangjun Song +3
Role-playing language agents (RPLAs) should play characters whose values and behavior evolve as the story progresses, not maintain a fixed persona. Existing benchmarks measure fact…
cs.CL2026
SimuHome: A Temporal- and Environment-Aware Benchmark for Smart Home LLM Agents
Gyuhyeon Seo, Jungwoo Yang, Junseong Pyo +3
We introduce , a high-fidelity smart home simulator and a benchmark of 600 episodes for LLM-based smart home agents. Existing smart home benchmarks treat the hom…
cs.CL2025
In-N-Out: A Parameter-Level API Graph Dataset for Tool Agents
Seungkyu Lee, Nalim Kim, Yohan Jo
Tool agents--LLM-based systems that interact with external APIs--offer a way to execute real-world tasks. However, as tasks become increasingly complex, these agents struggle to id…