3 papers
stat.AP2026
Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients
Michael Hardy, Joshua Gilbert, Benjamin Domingue
The validity of assessments, from large-scale AI benchmarks to human classrooms, depends on the quality of individual items, yet modern evaluation instruments often contain thousan…
stat.ME2024
Polytomous Explanatory Item Response Models for Item Discrimination: Assessing Negative-Framing Effects in Social-Emotional Learning Surveys
Joshua B. Gilbert, Lijin Zhang, Esther Ulitzsch +1
Modeling item parameters as a function of item characteristics has a long history but has generally focused on models for item location. Explanatory item response models for item d…
econ.EM2024
Estimating Heterogeneous Treatment Effects with Item-Level Outcome Data: Insights from Item Response Theory
Joshua B. Gilbert, Zachary Himmelsbach, James Soland +2
Analyses of heterogeneous treatment effects (HTE) are common in applied causal inference research. However, when outcomes are latent variables assessed via psychometric instruments…