Showing cs.AIShow all
3 papers · 1 filter
cs.AI2025
PillagerBench: Benchmarking LLM-Based Agents in Competitive Minecraft Team Environments
Olivier Schipper, Yudi Zhang, Yali Du +2
LLM-based agents have shown promise in various cooperative and strategic reasoning tasks, but their effectiveness in competitive multi-agent environments remains underexplored. To…
cs.AI2025
MEAL: A Benchmark for Continual Multi-Agent Reinforcement Learning
Tristan Tomilin, Luka van den Boogaard, Samuel Garcin +7
Benchmarks play a central role in reinforcement learning (RL) research, yet their computational constraints often shape what is studied. Despite the motivation of lifelong learning…
cs.AI2025
HASARD: A Benchmark for Vision-Based Safe Reinforcement Learning in Embodied Agents
Tristan Tomilin, Meng Fang, Mykola Pechenizkiy
Advancing safe autonomous systems through reinforcement learning (RL) requires robust benchmarks to evaluate performance, analyze methods, and assess agent competencies. Humans pri…