5 papers
ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib
Shane Caldwell
Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof…
ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents
Shane Caldwell, Max Harley, Ads Dawson +3
As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a…
Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors
Shane A. Caldwell, Moein Khazraee, Elena Agostini +27
We propose an architecture, called NVQLink, for connecting high-performance computing (HPC) resources to the control system of a quantum processing unit (QPU) to accelerate workloa…
PentestJudge: Judging Agent Behavior Against Operational Requirements
Shane Caldwell, Max Harley, Michael Kouremetis +2
We introduce PentestJudge, a system for evaluating the operations of penetration testing agents. PentestJudge is a large language model (LLM)-as-judge with access to tools that all…
AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models
Ads Dawson, Rob Mulla, Nick Landers +1
We introduce AIRTBench, an AI red teaming benchmark for evaluating language models' ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML…