Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation
Jiajun Jiang, Sharon Zheng, Natan Vidra +1
AI coding agent benchmarks rank agents with the Chen et al. (2021) pass@k estimator, but current implementations misapply it: they set n to the number of unit tests in a single sub…
cs.AI2026
OrchestraBench: Evaluating Multi-Agent Orchestration Failure Modes, Recovery, and Decomposition Quality
Yidian Chen, Yingzi Gu, Natan Vidra +2
Multi-agent orchestration frameworks are moving from demos to production, yet benchmarks typically report task accuracy without diagnosing why a pipeline failed, where a cascade be…
cs.AI2026
Posture and Sustainment Optimization Under Adversarial Uncertainty
Amelie Norris, Alyssa Lee, Natan Vidra +1
Pre-commitment posture, the assignment of military assets to theater locations before conflict scenarios resolve, is a critical and formally unsolved problem in joint operational p…