4 papers
Grounding Without Corrective Control: Truth-Tracking Profiles for Large Language Models
Brett Reynolds
Recent work suggests that some large language model representations have content or reference. Grounding can secure either without supplying live routes for correction. This paper…
When benchmark inferences do not compose: Projectibility in AI evaluation
Brett Reynolds
An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new ta…
Adversarial Pragmatics for AI Safety Evaluation: A Diagnostic Framework and Seed Benchmark for Language-Mediated Control
Brett Reynolds
Safety evaluations for language models increasingly depend on judgments about ambiguous natural-language behaviour: whether a model followed an instruction, refused appropriately,…
From Checklists to Clusters: A Homeostatic Account of AGI Evaluation
Brett Reynolds
Contemporary AGI evaluations report multidomain capability profiles, yet they typically assign symmetric weights and rely on snapshot scores. This creates two problems: (i) equal w…