PEN-STACK: A non-fabricating tool layer for language-model agents in genome writing
arXiv:2608.20412
Abstract
Background. Language-model agents are widely used in biology, but they report quantities without a verifiable source and pose unmanaged biosecurity risks. Genome writing sharpens both: a write plan must specify a location, writer enzyme, cargo, and delivery vehicle, all quantitative and interdependent, so without an integrated tool layer, the agent must supply them. We introduce PEN-STACK, an open tool layer that supplies them with guaranteed provenance. Results. PEN-STACK provides ten genome-writing design stages as twenty-two scope-aware tools, accessible via a software development kit, a Model Context Protocol server, and a Representational State Transfer interface, under a type-enforced invariant: every quantity must originate from a validated tool. Without tools, three model families fabricated 90.8% to 98.8% of the 240 required quantities under a naive prompt; coaching left a residual of 0 to 4, with no model certified at zero. Driving the tools, the same models fabricated nothing on a four-goal audit. A pre-emission biosecurity screen matched expert labels on all eight designs. The expression-robustness axis validated at exact-site resolution (ρ= 0.571, n = 1,506) but not at the coarser resolution served by default (ρ\approx 0.16), which returns a machine-readable downgrade flag. Eight of ten pre-registered claims did not pass, each flagged as machine-readable. Conclusions. On this evidence, grounding, not prompting or model scale, removes fabrication, and grounding requires a substrate; the grounded arm, a four-goal audit, warrants replication at the 240-field scale. PEN-STACK provides that substrate as open, importable code for agentic genome-engineering systems.