6 papers
Testing and Evaluation of Agentic AI Systems In Military Command and Control
Ulysse Richard, Heather Frase, Sarah Cao +3
Agentic AI systems are being procured for military command and control (C2) under public commitments to rigorous testing and human oversight. Whether such commitments can be discha…
Monitoring Agentic Systems Before They're Reliable
Marisa Ferrara Boston, Glen Hanson, Effi Georgala +2
Agentic systems entering production typically operate as partially integrated assemblies where structural defects, not task-level errors, dominate the failure landscape. At this ma…
Risk Management for Mitigating Benchmark Failure Modes: BenchRisk
Sean McGregor, Victor Lu, Vassil Tashev +8
Large language model (LLM) benchmarks inform LLM use decisions (e.g., "is this LLM safe to deploy for my use case and context?"). However, benchmarks may be rendered unreliable by…
Red Teaming for Generative AI, Report on a Copyright-Focused Exercise Completed in an Academic Medical Center
James Wen, Sahil Nalawade, Zhiwei Liang +38
Background: Generative artificial intelligence (AI) deployment in academic medical settings raises copyright compliance concerns. Dana-Farber Cancer Institute implemented GPT4DFCI,…
Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects
Reva Schwartz, Rumman Chowdhury, Akash Kundu +17
Conventional AI evaluation approaches concentrated within the AI stack exhibit systemic limitations for exploring, navigating and resolving the human and societal factors that play…
AILuminate: Introducing v1.0 of the AI Risk and Reliability Benchmark from MLCommons
Shaona Ghosh, Heather Frase, Adina Williams +99
The rapid advancement and deployment of AI systems have created an urgent need for standard safety-evaluation frameworks. This paper introduces AILuminate v1.0, the first comprehen…