Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Avyay M. Casheekar, Hariganesh Tangirala
Current agent evaluations score models on the state visible at the end of a stopped run which they count as one trial. However, interpreting the score as a final result would requi…
cs.AI2025
Adapting Probabilistic Risk Assessment for AI
Anna Katariina Wisakanto, Joe Rogero, Avyay M. Casheekar +1
Modern general-purpose artificial intelligence (AI) systems present an urgent risk management challenge, as their rapidly evolving capabilities and potential for catastrophic harm…