Bayesian Consensus Calibration of Continuously Evolving IRT Item Banks
arXiv:2609.13590
Abstract
AI-based item generation and NLP-based prediction of item parameters are producing item banks that are substantially larger, sparser, and more frequently updated than conventional banks. Hierarchical Bayesian item response theory (IRT) is a natural calibration framework for such banks, but the common practice of refitting the entire accumulated response history at each update is costly and can exceed available memory. We describe \emph{consensus calibration}, a divide-and-conquer procedure that calibrates each time period independently and reconstructs the pooled posterior in two layers. First, the posterior draws of each period are mapped to a common metric by a robust characteristic-curve linking (Haebara) that is solved separately for each draw, which propagates the uncertainty of the linking transformation into the linked posteriors. Second, the linked item posteriors are combined as a product of Gaussian densities from which the population prior contributed by each period is removed and a single prior---obtained by consensus across the per-period population posteriors---is reinstated. The correction targets the posterior dispersion, not only its location. As evidence for consensus calibration, we compare it to a pooled single-run analysis on a large operational assessment in terms of item-parameter recovery, an uncertainty-by-exposure diagnostic, and the ability distributions.
Accepted at AIME-Con 2026; to appear in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con), NCME