11 citations · 22 across the 8 of their papers we have counts for
3 papers · 1 filter
Testing and Evaluation of Agentic AI Systems In Military Command and Control
Ulysse Richard, Heather Frase, Sarah Cao +3
Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discha…
Monitoring Agentic Systems Before They're Reliable
Marisa Ferrara Boston, Glen Hanson, Effi Georgala +2
Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the failure landscape. At this ma…
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
Sean McGregor, Victor Lu, Vassil Tashev +8
Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by…