1 citations · 1 across the 2 of their papers we have counts for
Showing 2025Show all
2 papers · 1 filter
cs.CY2025
Who Should Run Advanced AI Evaluations -- AISIs?
Merlin Stein, Milan Gandhi, Theresa Kriecherbauer +2
Artificial Intelligence (AI) Safety Institutes and governments worldwide are deciding whether they evaluate advanced AI themselves, support a private evaluation ecosystem or do bot…
cs.AI2025
Measuring AI agent autonomy: Towards a scalable approach with code inspection
Peter Cihon, Merlin Stein, Gagan Bansal +2
AI agents are AI systems that can achieve complex goals autonomously. Assessing the level of agent autonomy is crucial for understanding both their potential benefits and risks. Cu…