2 papers
cs.CR2026
MOSAIC-Bench: Measuring Compositional Vulnerability Induction in Coding Agents
Jonathan Steinberg, Oren Gal
Coding agents often pass per-prompt safety review yet ship exploitable code when their tasks are decomposed into routine engineering tickets. The challenge is structural: existing…
cs.CR2026
Semantic Denial of Service in LLM-controlled robots
Jonathan Steinberg, Oren Gal
Safety-oriented instruction-following is supposed to keep LLM-controlled robots safe. We show it also creates an availability attack surface. By injecting short safety-plausible ph…