Towards AI epidemiology: a measurement standardisation framework for prospective risk detection
arXiv:2512.15783 · doi:10.1007/s00146-026-03276-3
Abstract
This paper proposes a measurement standardisation framework that compresses expert-AI interactions into structured, comparable fields for prospective risk detection in deployed AI systems, without access to model internals. This concept paper defines the framework's scope, semantically and statistically, and specifies a protocol for its empirical testing. The population-level claims it is designed to support therefore belong to a staged research programme rather than to results claimed here. Measurement standardisation underpins three claims. The first is a reliability claim: under bounded conditions, large language models can produce reliable, standardised assessments of the evidential and policy alignment of expert-AI interactions. The second is a governance claim: alignment scores give experts an immediate signal during deployment and give institutions a basis for monitoring alignment patterns across mission types, models, and domains. The third is an outcome validation claim: once measurement standardisation is established, aggregate alignment scores could be used to study associations with downstream outcomes in regulated professional settings. This introduces the possibility of an "AI epidemiology", a form of risk detection based on correlated variables instead of mechanistic analysis, inspired by epidemiological reasoning. A minimal application of the protocol to a published expert-AI corpus shows that the judge reproduces its policy and evidential alignment scores across two runs under the specified conditions. Judge reliability at scale remains to be validated in future work. The paper sets out a defined grammar of eight interaction fields, together with a statistical protocol based on paired bootstrap inference, DeLong's test for paired AUCs as a sensitivity check, a pre-specified one-sided non-inferiority margin of 0.05, and Holm-Bonferroni correction.
38 pages, 3 figures, 4 tables. Accepted for publication in AI & Society
References in corpus (9)
- Training language models to follow instructions with human feedback
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- LLM Evaluators Recognize and Favor Their Own Generations
- Sycophancy in Large Language Models: Causes and Mitigations
- The Butterfly Effect of Altering Prompts: How Small Changes and Jailbreaks Affect Large Language Model Performance
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- TokenSHAP: Interpreting Large Language Models with Monte Carlo Shapley Value Estimation
- Monitoring Machine Learning Systems: A Multivocal Literature Review
- No Free Labels: Limitations of LLM-as-a-Judge Without Human Grounding