4 papers
HarvestBench: Measuring Whether LLM Agents Will Pay to Avoid Killing Animals
Jasmine Brazilek, Miles Tidmarsh, Matthias Endres +2
HarvestBench is the first benchmark to 1) put a price on avoiding a side effect and 2) name the side effect as a living creature. Nine LLMs each drive a crew of two tractors to gat…
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
Jasmine Brazilek, Joel Christoph, Maheep Chaudhary +4
Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may…
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Jasmine Brazilek, Maheep Chaudhary, Zoe Lu +1
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…
Alignment midtraining for animals
Jasmine Brazilek, Miles Tidmarsh
We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal…