Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
Correct Looks Better: Pairwise Comparisons Reveal Accuracy Rankings
Mina Remeli, Moritz Hardt
Pairwise comparisons combined with aggregation methods like Elo have become central to evaluating generative models, yet concerns remain that they reward superficial stylistic cues…
cs.AI2026
Computational Arbitrage in AI Model Markets
Ricardo Olmedo, Bernhard Schölkopf, Moritz Hardt
Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a…