collaborators

5 papers

cs.LO2026

ProofJudge: Tool-Grounded LLM Evaluation of Formal Proof Quality in Mathlib

Shane Caldwell

Formal proofs in Lean 4 that pass the kernel's type checker can nonetheless vary widely in quality. We introduce ProofJudge, an agentic LLM-as-judge system that scores formal proof…

cs.CR2026

ScopeJudge: Cost-Aware Pre-Execution Gating for Offensive Security Agents

Shane Caldwell, Max Harley, Ads Dawson +3

As LLM agents take on offensive security work, a single out-of-scope tool call can breach a client's engagement boundary, disrupt production, or void a bug-bounty finding. Unlike a…

quant-ph2025

Platform Architecture for Tight Coupling of High-Performance Computing with Quantum Processors

Shane A. Caldwell, Moein Khazraee, Elena Agostini +27

We propose an architecture, called NVQLink, for connecting high-performance computing (HPC) resources to the control system of a quantum processing unit (QPU) to accelerate workloa…

cs.AI2025

PentestJudge: Judging Agent Behavior Against Operational Requirements

Shane Caldwell, Max Harley, Michael Kouremetis +2

We introduce PentestJudge, a system for evaluating the operations of penetration testing agents. PentestJudge is a large language model (LLM)-as-judge with access to tools that all…

cs.CR2025

AIRTBench: Measuring Autonomous AI Red Teaming Capabilities in Language Models

Ads Dawson, Rob Mulla, Nick Landers +1

We introduce AIRTBench, an AI red teaming benchmark for evaluating language models' ability to autonomously discover and exploit Artificial Intelligence and Machine Learning (AI/ML…