papers

Publications (13)

stat.ME2020

Estimating population average treatment effects from experiments with noncompliance

Kellie Ottoboni, Jason Poulos

Randomized control trials (RCTs) are the gold standard for estimating causal effects, but often use samples that are non-representative of the actual population of interest. We pro…

cs.SE2026

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Mike A. Merrill, Alexander G. Shaw, Nicholas Carlini +82

AI agents may soon become capable of autonomously completing valuable, long-horizon tasks in diverse domains. Current benchmarks either do not measure real-world tasks, or are not…

econ.GN2021

Amnesty Policy and Elite Persistence in the Postbellum South: Evidence from a Regression Discontinuity Design

Jason Poulos

This paper investigates the impact of Reconstruction-era amnesty policy on the officeholding and wealth of elites in the postbellum South. Amnesty policy restricted the political a…

cs.LG2026

Humanity's Last Exam

Long Phan, Alice Gatti, Ziwen Han +1144

Benchmarks are important tools for tracking the rapid advancements in large language model (LLM) capabilities. However, benchmarks are not keeping pace in difficulty: LLMs now achi…

cs.LG2022

Are deep learning models superior for missing data imputation in large surveys? Evidence from an empirical comparison

Zhenhua Wang, Olanrewaju Akande, Jason Poulos +1

Multiple imputation (MI) is a popular approach for dealing with missing data arising from non-response in sample surveys. Multiple imputation by chained equations (MICE) is one of…

stat.ML2018

Missing Data Imputation for Supervised Learning

Jason Poulos, Rafael Valle

Missing data imputation can help improve the performance of prediction models in situations where missing data hide useful information. This paper compares methods for imputing mis…

cs.AI2024

Adversarial Machine Learning: Bayesian Perspectives

David Rios Insua, Roi Naveiro, Victor Gallego +1

Adversarial Machine Learning (AML) is emerging as a major field aimed at protecting machine learning (ML) systems against security threats: in certain scenarios there may be advers…

stat.ME2021

Retrospective causal inference via matrix completion, with an evaluation of the effect of European integration on cross-border employment

Jason Poulos, Andrea Albanese, Andrea Mercatanti +1

We propose a method of retrospective counterfactual imputation in panel data settings with later-treated and always-treated units, but no never-treated units. We use the observed o…

cs.CV2021

Character-Based Handwritten Text Transcription with Attention Networks

Jason Poulos, Rafael Valle

The paper approaches the task of handwritten text recognition (HTR) with attentional encoder-decoder networks trained on sequences of characters, rather than words. We experiment o…

stat.AP2023

Targeted learning in observational studies with multi-valued treatments: An evaluation of antipsychotic drug treatment safety

Jason Poulos, Marcela Horvitz-Lennon, Katya Zelevinsky +6

We investigate estimation of causal effects of multiple competing (multi-valued) treatments in the absence of randomization. Our work is motivated by an intention-to-treat study of…

econ.GN2023

Gender gaps in frontier entrepreneurship? Evidence from 1901 Oklahoma land lottery winners

Jason Poulos

The paper investigates gender differences in entrepreneurship by exploiting a large-scale land lottery in Oklahoma at the turn of the 20 century. Lottery winners clai…

econ.GN2023

State-Building through Public Land Disposal? An Application of Matrix Completion for Counterfactual Prediction

Jason Poulos

This paper examines how homestead policies, which opened vast frontier lands for settlement, influenced the development of American frontier states. It uses a treatment propensity-…

stat.ML2021

RNN-based counterfactual prediction, with an application to homestead policy and public schooling

Jason Poulos, Shuxi Zeng

This paper proposes a method for estimating the effect of a policy intervention on an outcome over time. We train recurrent neural networks (RNNs) on the history of control unit ou…