Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
Miles Q. Li, Benjamin C. M. Fung, Martin Weiss +3
As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Cur…
cs.AI2026
Adaptive Prompt Embedding Optimization for LLM Jailbreaking
Miles Q. Li, Benjamin C. M. Fung, Boyang Li +2
Existing white-box jailbreak attacks against aligned LLMs typically append discrete adversarial suffixes to the user prompt, which visibly alters the prompt and operates in a combi…