From the 1 of 9 linked papers with an AI index.
4 papers · 1 filter
From Sycophantic Consensus to Pluralistic Repair: Why AI Alignment Must Surface Disagreement
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka
Pluralistic alignment is typically operationalised as preference aggregation: producing responses that span (Overton), steer toward (Steerable), or proportionally represent (Distri…
The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Recent published evidence from frontier laboratories shows that contemporary AI models can recognise evaluation contexts, latently represent them, and behave differently under thos…
Deployment-Relevant Alignment Cannot Be Inferred from Model-Level Evaluation Alone
Varad Vishwarupe, Nigel Shadbolt, Marina Jirotka +1
Alignment evaluation in machine learning has largely become evaluation of models. Influential benchmarks score model outputs under fixed inputs, such as truthfulness, instruction f…
A Scalable Communication Protocol for Networks of Large Language Models
Samuele Marro, Emanuele La Malfa, Jesse Wright +4
Communication is a prerequisite for collaboration. When scaling networks of AI-powered agents, communication must be versatile, efficient, and portable. These requisites, which we…