works on

From the 1 of 9 linked papers with an AI index.

activity
20242026
collaborators

9 papers

cs.LG2026

Toward Human Rights Benchmarking for LLMs: A Pilot Methodology

Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee +3

Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can…

cs.CY2026

L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education

James Edgell, Wm. Matthew Kennedy, Ben Knight +3

The paper presents L2-Bench, an open‑source benchmark of over 1,000 task‑response pairs designed to evaluate large language models on competencies relevant to second language (L2)…

cs.AI2026

Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results

Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45

AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…

cs.AI2026

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting

Avijit Ghosh, Anka Reuel, Jenny Chim +45

AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…

cs.HC2026

Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems

Ben Knight, Wm. Matthew Kennedy, Danielle Carvalho +2

AI-powered language learning tools increasingly provide instant, personalised feedback to millions of learners worldwide. However, this feedback can fail in ways that are difficult…

cs.CY2026

Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench

James Edgell, Wm. Matthew Kennedy, Isaac Pattis +3

The rapid adoption of large language models in AI-powered language education has created an urgent need for evaluations that assess pedagogical effectiveness, particularly in langu…