3 papers
cs.CY2026
Frontier Lag: A Bibliometric Audit of Capability Misrepresentation in Academic AI Evaluation
David Gringras, Misha Salahshoor
Readers of applied-domain LLM capability evaluations want to know what AI systems can currently do. That literature answers a related, but consequentially different, question: what…
cs.AI2026
IatroBench: Pre-Registered Evidence of Iatrogenic Harm from AI Safety Measures
David Gringras
A heavily safety-trained model will hand a physician the full, patient-followable benzodiazepine taper and refuse it to the patient who needs it, over identical clinical facts; the…
cs.AI2026
Safety Under Scaffolding: How Evaluation Conditions Shape Measured Safety
David Gringras
A safety score earned on a benchmark need not predict how the same model behaves once it is wrapped in an agentic scaffold the benchmark never tested. We ran six frontier models th…