Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
ClawsBench: Evaluating Capability and Safety of LLM Productivity Agents in Simulated Workspaces
Xiangyi Li, Kyoung Whan Choe, Yimin Liu +12
Large language model (LLM) agents are increasingly deployed to automate productivity tasks (e.g., email, scheduling, document management), but evaluating them on live services is r…
cs.AI2024
Massively Multiagent Minigames for Training Generalist Agents
Kyoung Whan Choe, Ryan Sullivan, Joseph Suárez
We present Meta MMO, a collection of many-agent minigames for use as a reinforcement learning benchmark. Meta MMO is built on top of Neural MMO, a massively multiagent environment…