2 papers
cs.AI2026
DSGBench: A Diverse Strategic Game Benchmark for Evaluating LLM-based Agents in Complex Decision-Making Environments
Wenjie Tang, Yuan Zhou, Erqiang Xu +3
Large language model (LLM)-based agents are increasingly applied to complex strategic environments that demand long-horizon reasoning, multi-agent interaction, and decision-making…
cs.CR2026
When Convenience Becomes Risk: A Semantic View of Under-Specification in Host-Acting Agents
Di Lu, Yongzhi Liao, Xutong Mu +5
Host-acting agents promise a convenient interaction model in which users specify goals and the system determines how to realize them. We argue that this convenience introduces a di…