paper

BioFirewall: A genome-writing-native governance layer for design-stage biosecurity screening of agentic AI

arXiv:2608.20413

Abstract

Background. Artificial-intelligence design tools now plan genome-scale edits, and agentic systems execute those plans with progressively less human oversight. Biosecurity controls are limited to two points: refusal guardrails at the foundation model and sequence-identity screening at the synthesiser. The design stage between them, where the plan is specified, remains governed by recommendations rather than any deployed system. Results. We present BioFirewall, a rule-governed middleware that intercepts a genome-writing plan and returns allow, flag-for-review, or refuse across five hazard axes native to genome writing: cargo, locus, edit type, germline and scale, with cited evidence, a signed design passport, a tamper-evident audit log, and tiered access. On a de-circularised benchmark of safe proxies scored against independent oracles, a function-aware cargo classifier reached a true-positive rate of 0.72 (95% CI 0.43 to 0.89) at a 1% false-positive rate, whereas frontier and open language-model judges did not screen the same sequences reliably. Under prompt injection, the open-weight judges flipped their blocking verdict to allow in 3 and 5 of 6 trials per channel, while the deterministic screen remained invariant. None of 288 legitimate plans from three templates was refused, yielding a certified 95% upper bound of 0.0103 on the false-refuse rate, and a session monitor intercepted cross-call decomposition attacks. On a held-out gene set, the locus axis was enriched for drivers of in vivo insertional oncogenesis (AUROC 0.605; odds ratio 3.34). Conclusions. Design-stage governance is achievable in practice. BioFirewall is released as open source with a pre-registered, open-data-reproducible benchmark.

BioFirewall: A genome-writing-native governance layer for design-stage biosecurity screening of agentic AI · wovepaper