3 papers
cs.CR2026
MonitoringBench: Semi-Automated Red-Teaming for Agent Monitoring
Monika JotautaitÄ, Maria Angelica Martinez, Ollie Matthews +1
We introduce a red-teaming methodology that exposes harder-to-catch attacks for coding-agent monitors, suggesting that current practices may under-elicit attacks and overstate moni…
cs.CL2025
Speciesism in AI: Evaluating Discrimination Against Animals in Large Language Models
Monika JotautaitÄ, Lucius Caviola, David A. Brewster +1
As large language models (LLMs) become more widely deployed, it is crucial to examine their ethical tendencies. Building on research on fairness and discrimination in AI, we invest…
cs.CY2025
From Stability to Inconsistency: A Study of Moral Preferences in LLMs
Monika Jotautaite, Mary Phuong, Chatrik Singh Mangat +1
As large language models (LLMs) increasingly integrate into our daily lives, it becomes crucial to understand their implicit biases and moral tendencies. To address this, we introd…