3 papers
cs.CY2025
Deprecating Benchmarks: Criteria and Framework
Ayrton San Joaquin, Rokas Gipiškis, Leon Staufer +1
As frontier artificial intelligence (AI) models rapidly advance, benchmarks are integral to comparing different models and measuring their progress in different task-specific domai…
cs.CY2025
Mapping Industry Practices to the EU AI Act's GPAI Code of Practice Safety and Security Measures
Lily Stelling, Mick Yang, Rokas Gipiškis +5
This report provides a detailed comparison between the Safety and Security measures proposed in the EU AI Act's General-Purpose AI (GPAI) Code of Practice (Third Draft) and the cur…
cs.CY2025
Audit Cards: Contextualizing AI Evaluations
Leon Staufer, Mick Yang, Anka Reuel +1
AI governance frameworks increasingly rely on audits, yet the results of their underlying evaluations require interpretation and context to be meaningfully informative. Even techni…