3 papers
cs.AI2026
From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution
Stefan G. Creadore
Large-language-model agents can propose and execute actions, but proposal, authority, dispatch, verified external effect, and serving promotion are different claims. We present Pra…
cs.SE2026
When Agent Metrics Measure Different Things: An Evidence-Grounded Audit of the Praxa AI Pipeline
Stefan G. Creadore, Peyton Woakz
Agent evaluations can be numerically correct while measuring a different construct from the one implied by their labels. We present a retrospective measurement audit of selected Pr…
cs.AI2026
Plato-Bio: verification-first biological novelty screening with temporal rediscovery and structural benchmarks
Stefan G. Creadore
Large language model research agents can connect literature retrieval, analysis code, and manuscript preparation, but coherent output does not establish scientific validity. We dev…