10 papers
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
Mechanistic Interpretability Needs Philosophy
Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6
Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…
Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing
Ruchira Dhar, Anders Søgaard
Recent advances in large language models (LLMs) have prompted a growing body of work that questions the methodology of prevailing evaluation practices. However, many such critiques…
Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives
Ruchira Dhar, Qiwei Peng, Anders Søgaard
Compositionality is considered central to language abilities. As performant language systems, how do large language models (LLMs) do on compositional tasks? We evaluate adjective-n…
Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research
Ninell Oldenburg, Ruchira Dhar, Anders Søgaard
In this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelli…
On the Measure of a Model: From Intelligence to Generality
Ruchira Dhar, Ninell Oldenburg, Anders Soegaard
Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence…