11 citations · 12 across the 5 of their papers we have counts for
11 papers
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82
AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…
Gender gaps in frontier entrepreneurship? Evidence from 1901 Oklahoma land lottery winners
Jason Poulos
The paper investigates gender differences in entrepreneurship by exploiting a large-scale land lottery in Oklahoma at the turn of the 20 century. Lottery winners clai…
Retrospective causal inference via matrix completion, with an evaluation of the effect of European integration on cross-border employment
Jason Poulos, Andrea Albanese, Andrea Mercatanti +1
We propose a method of retrospective counterfactual imputation in panel data settings with later-treated and always-treated units, but no never-treated units. We use the observed o…
Amnesty Policy and Elite Persistence in the Postbellum South: Evidence from a Regression Discontinuity Design
Jason Poulos
This paper investigates the impact of Reconstruction-era amnesty policy on the officeholding and wealth of elites in the postbellum South. Amnesty policy restricted the political a…
Are deep learning models superior for missing data imputation in large surveys? Evidence from an empirical comparison
Zhenhua Wang, Olanrewaju Akande, Jason Poulos +1
Multiple imputation (MI) is a popular approach for dealing with missing data arising from non-response in sample surveys. Multiple imputation by chained equations (MICE) is one of…
Adversarial Machine Learning: Bayesian Perspectives
David Rios Insua, Roi Naveiro, Victor Gallego +1
Adversarial Machine Learning (AML) is emerging as a major field aimed at protecting machine learning (ML) systems against security threats: in certain scenarios there may be advers…