3 papers
cs.CR2026
Evaluating Frontier AI Agents as Autonomous Clinical Security Auditors
Michael O. Eniolade
Clinical AI models can expose patients to harm when adversarial vulnerabilities go undetected, yet formal security auditing requires statistical expertise, specialized tools, and s…
cs.LG2026
Calibration, Uncertainty Communication, and Deployment Readiness in CKD Risk Prediction: A Framework Evaluation Study
Michael O. Eniolade
Machine learning models for chronic kidney disease (CKD) risk prediction often post strong discrimination scores on internal test sets. Calibration and uncertainty quantification g…
cs.LG2026
StepShield: When, Not Whether to Intervene on Rogue Agents
Gloria Felicia, Zitha Sasindran, Jinfeng He +3
Agent safety benchmarks measure whether a monitor detects harm, not when. Yet timing is the difference between intervention and autopsy. We introduce StepShield, the first benchmar…