From the 1 of 21 linked papers with an AI index.
21 papers
Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory
Jon-Paul Cacioli
The paper applies signal detection theory to separate a language model’s factual accuracy from the quality of its confidence estimates, introducing a model‑free metric (meta‑I_2r)…
Beyond the Mean: Within-Model Reliable Change Detection for LLM Evaluation
Jon-Paul Cacioli
We adapted the Reliable Change Index (RCI; Jacobson and Truax, 1991) from clinical psychology to item-level LLM version comparison on 2,000 MMLU-Pro items (K=10 samples at T=0.7).…
Instruction Complexity Induces Positional Collapse in Adversarial LLM Evaluation
Jon-Paul Cacioli
When instructed to underperform on multiple-choice evaluations, do language models engage with question content or fall back on positional shortcuts? We map the boundary between th…
Option-Order Randomisation Reveals a Distributional Position Attractor in Prompted Sandbagging
Jon-Paul Cacioli
A predecessor pilot (Cacioli, 2026) found that Llama-3-8B implements prompted sandbagging as positional collapse rather than answer avoidance. However, fixed option ordering in MML…
Below-Chance Blindness: Prompted Underperformance in Small LLMs Produces Positional Bias Rather than Answer Avoidance
Jon-Paul Cacioli
Detecting sandbagging--the deliberate underperformance on capability evaluations--is an open problem in AI safety. We tested whether symptom validity testing (SVT) logic from clini…
Distilling Self-Consistency into Verbal Confidence: A Pre-Registered Negative Result and Post-Hoc Rescue on Gemma 3 4B
Jon-Paul Cacioli
Small instruct-tuned LLMs produce degenerate verbal confidence under minimal elicitation: ceiling rates above 95%, near-chance Type-2 AUROC, and Invalid validity profiles. We test…