From the 1 of 3 linked papers with an AI index.
3 papers
cs.AI2026
From Controlled to the Wild: Evaluation of Pentesting Agents for the Real-World
Pedro Conde, Henrique Branquinho, Valerio Mazzone +3
The paper introduces a practical evaluation protocol for AI-driven penetration testing agents that emphasizes validated vulnerability discovery in complex, realistic targets, and p…
cs.LG2026
Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents
Yujun Zhou, Kehan Guo, Haomin Zhuang +8
Interactive LLM agents are becoming part of daily work, but they do not reliably become easier to work with over time: a correction remembered in one session may still be violated…
cs.AI2026
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…