6 papers
Coercion and Deception in AI-to-AI Management: An Agentic Benchmark of Unprompted Escalation
Jasmine Brazilek, Maheep Chaudhary, Zoe Lu +1
Multi-agent systems routinely place one AI agent in authority over another. When a subordinate refuses a task, the manager chooses the outcome: it can renegotiate, report the failu…
Assert, don't describe: Linguistic features that shift LLM reasoning about animal welfare
Jasmine Brazilek, Harper Dunn
Animal-welfare advocates produce a lot of writing, and increasingly that writing trains the language models that millions of people then ask about animal welfare. Using vocabulary-…
Helpfulness Hurts: Domain-Dependent Degradation of Mid-Trained Compassion Values Under Post-Training
Jasmine Brazilek, Juliana Seawell
Standard post-training pipelines apply supervised fine-tuning (SFT) and reinforcement learning (RL) to make language models helpful, but these processes may inadvertently degrade v…
Small edits, large models: How Wikipedia advocacy shapes LLM values
Jasmine Brazilek, Maria Navas, Alexa Gnauck
Can a small group of volunteers shape how AI systems discuss animal welfare, just by editing Wikipedia? We show that they can. Wikipedia appears in nearly every major language mode…
Your AI Travel Agent Would Book You a Bullfight: An Agentic Benchmark for Implicit Animal Welfare in Frontier AI Models
Jasmine Brazilek, Joel Christoph, Maheep Chaudhary +4
Previous research has evaluated animal welfare using question-and-answer benchmarks. This study investigates whether these evaluations also hold in agentic settings. The agents may…
Alignment midtraining for animals
Jasmine Brazilek, Miles Tidmarsh
We investigate the robustness of value alignment via midtraining with synthetic documents, using animal compassion as a value that is both important in its own right and orthogonal…