Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka
Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distri…
cs.AI2026
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and behave differently under thos…
cs.AI2026
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction f…