3 papers
cs.AI2026
When Is an Agent Evaluation Over? Outcome Finality and Cross-Unit Separation
Avyay M. Casheekar, Hariganesh Tangirala
Agent evaluations commonly score the state observed when a run stops and count the run as one trial. Interpreting that score as a final result from a separate trial requires outcom…
cs.CY2026
An Evaluation Framework for National AI Regulation
Kaushik Sanjay Prabhakar, Tarun Adarsh R S, Amal Dhivyan Gregory +3
Governments use laws, institutions, funding programs and nonbinding guidance to shape how AI is developed and used. Comparing these national approaches is difficult. A binding rule…
cs.AI2025
Adapting Probabilistic Risk Assessment for AI
Anna Katariina Wisakanto, Joe Rogero, Avyay M. Casheekar +1
Modern general-purpose artificial intelligence (AI) systems present an urgent risk management challenge, as their rapidly evolving capabilities and potential for catastrophic harm…