1 citations · 1 across the 2 of their papers we have counts for
2 papers
cs.AI2026
CivBench: A Long-Horizon Benchmark for Tool-Mediated Agents in Civilization VI
Austin Tudor David Andrews, Liam Wilkinson, Jamie Heagerty +3
We present CivBench, an open-source benchmark for evaluating language model agents in long-horizon, tool-mediated environments through the Model Context Protocol (MCP). A single ep…
cs.AI2025★ 1 cited
A Real-World Evaluation of LLM Medication Safety Reviews in NHS Primary Care
Oliver Normand, Esther Borsi, Mitch Fruin +6
Large language models (LLMs) often match or exceed clinician-level performance on medical benchmarks, yet very few are evaluated on real clinical data or examined beyond headline m…