conformal inference 1extractable memorization 1large language models 1privacy evaluation 1statistical testing 1
From the 1 of 8 linked papers with an AI index.
2 citations · 2 across the 2 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
BountyBench: Dollar Impact of AI Agent Attackers and Defenders on Real-World Cybersecurity Systems
Andy K. Zhang, Joey Ji, Celeste Menders +31
AI agents have the potential to significantly alter the cybersecurity landscape. Here, we introduce the first framework to capture offensive and defensive cyber-capabilities in evo…
cs.CR2025
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
Andy K. Zhang, Neil Perry, Riya Dulepet +24
Language Model (LM) agents for cybersecurity that are capable of autonomously identifying vulnerabilities and executing exploits have potential to cause real-world impact. Policyma…