2 papers
cs.CR2026
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering
Mary Llewellyn, Isobel Thornton, James Bishop +1
LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations…
stat.CO2025
Grid Particle Gibbs with Ancestor Sampling for State-Space Models
Mary Llewellyn, Ruth King, VÃctor Elvira +1
We consider the challenge of estimating the model parameters and latent states of general state-space models within a Bayesian framework. We extend the commonly applied particle Gi…