From the 1 of 3 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
Pedro Conde, Henrique Branquinho, Valerio Mazzone +3
The paper introduces a practical evaluation protocol for AI-driven penetration testing agents that emphasizes validated vulnerability discovery in complex, realistic targets, and p…
cs.AI2026
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…