activity
20242026
collaborators

10 papers

cs.AI2026

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Avijit Ghosh, Anka Reuel, Jenny Chim +45

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…

cs.CL2026

Mechanistic Interpretability Needs Philosophy

Iwan Williams, Ninell Oldenburg, Ruchira Dhar +6

Mechanistic interpretability (MI) aims to explain how neural networks work by uncovering their underlying mechanisms. As the field grows in influence, it is increasingly important…

cs.CL2026

Evaluation Revisited: A Taxonomy of Evaluation Concerns in Natural Language Processing

Ruchira Dhar, Anders Søgaard

Recent advances in large language models (LLMs) have prompted a growing body of work that questions the methodology of prevailing evaluation practices. However, many such critiques…

cs.CL2026

Evaluating Adjective-Noun Compositionality in LLMs: Functional vs Representational Perspectives

Ruchira Dhar, Qiwei Peng, Anders Søgaard

Compositionality is considered central to language abilities. As performant language systems, how do large language models (LLMs) do on compositional tasks? We evaluate adjective-n…

cs.AI2025

Realist and Pluralist Conceptions of Intelligence and Their Implications on AI Research

Ninell Oldenburg, Ruchira Dhar, Anders Søgaard

In this paper, we argue that current AI research operates on a spectrum between two different underlying conceptions of intelligence: Intelligence Realism, which holds that intelli…

cs.AI2025

On the Measure of a Model: From Intelligence to Generality

Ruchira Dhar, Ninell Oldenburg, Anders Soegaard

Benchmarks such as ARC, Raven-inspired tests, and the Blackbird Task are widely used to evaluate the intelligence of large language models (LLMs). Yet, the concept of intelligence…