2 papers
cs.CR2026
SIR-Bench: Evaluating Investigation Depth in Security Incident Response Agents
Daniel Begimher, Cristian Leo, Jack Huang +2
We present SIR-Bench, a benchmark of 794 test cases for evaluating autonomous security incident response agents that distinguishes genuine forensic investigation from alert parroti…
cs.RO2026
Alpamayo-R1: Bridging Reasoning and Action Prediction for Generalizable Autonomous Driving in the Long Tail
NVIDIA, :, Yan Wang +41
End-to-end architectures trained via imitation learning have advanced autonomous driving by scaling model size and data, yet performance remains brittle in safety-critical long-tai…