Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Avyay M. Casheekar, Hariganesh Tangirala
Agent evaluations commonly score the state observed when a run stops and count the run as one trial. Interpreting that score as a final result from a separate trial requires outcom…
cs.AI2025
Adapting Probabilistic Risk Assessment for AI
Anna Katariina Wisakanto, Joe Rogero, Avyay M. Casheekar +1
Modern general-purpose artificial intelligence (AI) systems present an urgent risk management challenge, as their rapidly evolving capabilities and potential for catastrophic harm…