Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
I must delete the evidence: AI Agents Explicitly Cover up Fraud and Violent Crime
Thomas Rivasseau
As ongoing research explores the ability of AI agents to be insider threats and act against company interests, we showcase the abilities of such agents to act against human well be…
cs.AI2025
Invasive Context Engineering to Control Large Language Models
Thomas Rivasseau
Current research on operator control of Large Language Models improves model robustness against adversarial attacks and misbehavior by training on preference examples, prompting, a…