10 papers
Measuring Safety Alignment Effects in Autonomous Security Agents
Isaac David, Arthur Gervais
Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Single-turn refusal benchmarks ca…
Benchmarking Mythos-Linked Bug Rediscovery
Isaac David, Arthur Gervais
Anthropic's April 2026 Mythos materials combine benchmark claims with concrete bug-finding stories across OpenBSD, FreeBSD, Linux, FFmpeg, and browsers. This paper reports a contro…
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
Isaac David, Arthur Gervais
Safety-aligned language models often refuse cybersecurity requests whose wording resembles misuse, even when the task is authorized and defensive. This makes security evaluation am…
CrackMeBench: Binary Reverse Engineering for Agents
Isaac David, Arthur Gervais
Benchmarks for coding agents increasingly measure source-level software repair, and cybersecurity benchmarks increasingly measure broad capture-the-flag performance. Classical bina…
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
Isaac David, Arthur Gervais
Security updates create a short but important window in which defenders and attackers can compare vulnerable and patched software. Yet in many operational settings, the most access…
Alignment Contracts for Agentic Security Systems
Isaac David, Marco Guarnieri, Arthur Gervais
Agentic security systems increasingly combine LLM planners with tools that can discover, validate, and report vulnerabilities. This creates an asymmetric control problem: the syste…