Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Testing Interchangeability in LLM Agent Teams
Jianxin Gao, Tianyi Yu, Linna Deng +2
Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that a…
cs.AI2026
MirrorCraft: Paired Evaluation under Hidden Rule Changes in Minecraft
Jianxin Gao, Beini Hu, Runze Li +6
With the prosperity of the large language models (LLMs), it has become an interesting topic: how do LLM-based agents work in Minecraft? Unfortunately, most existing benchmarks eval…