1 citations · 1 across the 8 of their papers we have counts for
3 papers · 1 filter
Rethinking Uncertainty Evaluation in Large Language Models
Krish Matta, Atharv Naphade, Andy Zou
Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and d…
Comparing AI Agents to Cybersecurity Professionals in Real-World Penetration Testing
Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper +10
We present the first comprehensive evaluation of AI agents against human cybersecurity professionals in a live enterprise environment. We evaluate ten cybersecurity professionals a…
A Definition of AGI
Dan Hendrycks, Dawn Song, Christian Szegedy +30
The lack of a concrete definition for Artificial General Intelligence (AGI) obscures the gap between today's specialized AI and human-level cognition. This paper introduces a quant…