8 papers · 1 filter
Measuring Safety Alignment Effects in Autonomous Security Agents
Isaac David, Arthur Gervais
Do stock safety-aligned language models and their uncensored or abliterated derivatives behave differently when run as autonomous security agents? Single-turn refusal benchmarks ca…
Ablating Safety: Mechanisms for Removing Alignment in Language Models for Security Applications
Isaac David, Arthur Gervais
Safety-aligned language models often refuse cybersecurity requests whose wording resembles misuse, even when the task is authorized and defensive. This makes security evaluation am…
Patch2Vuln: Agentic Reconstruction of Vulnerabilities from Linux Distribution Binary Patches
Isaac David, Arthur Gervais
Security updates create a short but important window in which defenders and attackers can compare vulnerable and patched software. Yet in many operational settings, the most access…
Alignment Contracts for Agentic Security Systems
Isaac David, Marco Guarnieri, Arthur Gervais
Agentic security systems increasingly combine LLM planners with tools that can discover, validate, and report vulnerabilities. This creates an asymmetric control problem: the syste…
Towards Optimal Agentic Architectures for Offensive Security Tasks
Isaac David, Arthur Gervais
Agentic security systems increasingly audit live targets with tool-using LLMs, but prior systems fix a single coordination topology, leaving unclear when additional agents help and…
Multi-Agent Penetration Testing AI for the Web
Isaac David, Arthur Gervais
AI-powered development platforms are making software creation accessible to a broader audience, but this democratization has triggered a scalability crisis in security auditing. Wi…