3 papers
cs.AI2026
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka
Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distri…
cs.AI2026
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and behave differently under thos…
cs.CY2026
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Frontier AI safety claims - published assertions that a highly capable general-purpose model is below a threshold of concern, adequately mitigated, or suitable for release - increa…