1 paper
Quinn Dougherty, Max von Hippel, Simon Henniger +2
We present a benchmark for evaluating AI models and agents on real-world formal software verification tasks. We first scrape 11,039 property-based tests (PBTs) from real-world Pyth…