Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
cs.AI2026
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
Pedro Conde, Henrique Branquinho, Valerio Mazzone +3
AI pentesting agents are increasingly credible as offensive security systems, but current benchmarks still provide limited guidance on which will perform best in real-world targets…