collaborators

6 papers

cs.AI2026

AI Cartography: Mapping the Latent Landscape of AI Benchmark Ecosystems

Michael Hardy, Anka Reuel, Lijin Zhang +6

While aggregate leaderboard scores drive AI development, they contain substantial measurement noise whose sources and magnitudes remain unquantified, making it unclear when ranking…

cs.CY2026

Understanding Student Effort Using Response-Time Propensities During Problem Solving

Conrad Borchers, Lijin Zhang, Kexin Yang +2

Adaptive learning systems can produce substantial learning gains, yet many students engage for too brief or too superficial a period to benefit. A central obstacle is measuring eff…

stat.AP2026

Efficient Detection of Bad Benchmark Items with Novel Scalability Coefficients

Michael Hardy, Joshua Gilbert, Benjamin Domingue

The validity of assessments, from large-scale AI benchmarks to human classrooms, depends on the quality of individual items, yet modern evaluation instruments often contain thousan…

cs.AI2025

Fantastic Bugs and Where to Find Them in AI Benchmarks

Sang Truong, Yuheng Tu, Michael Hardy +8

Benchmarks are pivotal in driving AI progress, and invalid benchmark questions frequently undermine their reliability. Manually identifying and correcting errors among thousands of…

cs.CY2025

Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study

Calvin Isley, Joshua Gilbert, Evangelos Kassos +9

While large language models (LLMs) challenge conventional methods of teaching and learning, they present an exciting opportunity to improve efficiency and scale high-quality instru…

cs.CY2025

Measurement to Meaning: A Validity-Centered Framework for AI Evaluation

Olawale Salaudeen, Anka Reuel, Ahmed Ahmed +6

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning ca…