3 papers
stat.AP2026
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
Michael Hardy, Joshua Gilbert, Benjamin Domingue
The validity of assessments, from large-scale AI benchmarks to human classrooms, depends on the quality of individual items, yet modern evaluation instruments often contain thousan…
econ.EM2025
Estimating Heterogeneous Treatment Effects with Item-Level Outcome Data: Insights from Item Response Theory
Joshua B. Gilbert, Zachary Himmelsbach, James Soland +2
Analyses of heterogeneous treatment effects (HTE) are common in applied causal inference research. However, when outcomes are latent variables assessed via psychometric instruments…
stat.ME2025
Polytomous Explanatory Item Response Models for Item Discrimination: Assessing Negative-Framing Effects in Social-Emotional Learning Surveys
Joshua B. Gilbert, Lijin Zhang, Esther Ulitzsch +1
Modeling item parameters as a function of item characteristics has a long history but has generally focused on models for item location. Explanatory item response models for item d…