1 paper
Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty +3
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single ep…