1 citations · 1 across the 5 of their papers we have counts for
Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
From Human-Level AI Tales to AI Leveling Human Scales
Peter Romero, Fernando MartÃnez-Plumed, Zachary R. Tidler +11
Comparing AI models to "human level" is often misleading when benchmark scores are incommensurate or human baselines are drawn from a narrow population. To address this, we propose…
cs.LG2025
What should an AI assessor optimise for?
Daniel Romero-Alvarado, Fernando MartÃnez-Plumed, José Hernández-Orallo
An AI assessor is an external, ideally indepen-dent system that predicts an indicator, e.g., a loss value, of another AI system. Assessors can lever-age information from the test r…