Showing cs.AIShow all
3 papers · 1 filter
cs.AI2026
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
cs.AI2025
Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
Ninell Oldenburg, Ruchira Dhar, Anders Søgaard
In this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelli…
cs.AI2025
On the Measure of a Model: From Intelligence to Generality
Ruchira Dhar, Ninell Oldenburg, Anders Soegaard
Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence…