From the 1 of 9 linked papers with an AI index.
9 papers
Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
Savannah Thais, Wm. Matthew Kennedy, Abhigyan Acherjee +3
Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can…
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight +3
The paper presents L2-Bench, an open‑source benchmark of over 1,000 task‑response pairs designed to evaluate large language models on competencies relevant to second language (L2)…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting
Avijit Ghosh, Anka Reuel, Jenny Chim +45
AI evaluation results are produced at scale but reported inconsistently across leaderboards, model cards, benchmark papers, and company blogs. The cost is interpretive: readers can…
Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems
Ben Knight, Wm. Matthew Kennedy, Danielle Carvalho +2
AI-powered language learning tools increasingly provide instant, personalised feedback to millions of learners worldwide. However, this feedback can fail in ways that are difficult…
Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench
James Edgell, Wm. Matthew Kennedy, Isaac Pattis +3
The rapid adoption of large language models in AI-powered language education has created an urgent need for evaluations that assess pedagogical effectiveness, particularly in langu…