4 papers
Logical Consistency Between Disagreeing Experts and Its Role in AI Safety
Andrés Corrada-Emmanuel
If two experts disagree on a test, we may conclude both cannot be 100 per cent correct. But if they completely agree, no possible evaluation can be excluded. This asymmetry in the…
No-Knowledge Alarms for Misaligned LLMs-as-Judges
Andrés Corrada-Emmanuel
If we use LLMs as judges to evaluate the complex decisions of other LLMs, who or what monitors the judges? Infinite monitoring chains are inevitable whenever we do not know the gro…
Algebraic Evaluation Theorems
Andrés Corrada-Emmanuel
Majority voting (MV) is the prototypical ``wisdom of the crowd'' algorithm. Theorems considering when MV is optimal for group decisions date back to Condorcet's 1785 jury \emph{dec…
A logical alarm for misaligned binary classifiers
Andrés Corrada-Emmanuel, Ilya Parker, Ramesh Bharadwaj
If two agents disagree in their decisions, we may suspect they are not both correct. This intuition is formalized for evaluating agents that have carried out a binary classificatio…