Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Jasmine Brazilek, Miles Tidmarsh, Matthias Endres +2
HarvestBench is the first benchmark to 1) put a price on avoiding a side effect and 2) name the side effect as a living creature. Nine LLMs each drive a crew of two tractors to gat…
cs.AI2026
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
Jasmine Brazilek, Joel Christoph, Maheep Chaudhary +4
Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may…