2 papers
cs.CR2026
Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering
Mary Llewellyn, Isobel Thornton, James Bishop +1
LLM benchmarking metrics often misstate performance and uncertainty as they rely on two assumptions that frequently do not hold in practice: (i) a sufficient number of evaluations…
stat.ME2025
Statistical exploration of the Manifold Hypothesis
Nick Whiteley, Annie Gray, Patrick Rubin-Delanchy
The Manifold Hypothesis is a widely accepted tenet of Machine Learning which asserts that nominally high-dimensional data are in fact concentrated near a low-dimensional manifold,…