From the 1 of 4 linked papers with an AI index.
4 papers
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight +3
The paper presents L2-Bench, an open‑source benchmark of over 1,000 task‑response pairs designed to evaluate large language models on competencies relevant to second language (L2)…
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…
Ceci n'est pas une explication: Evaluating Explanation Failures as Explainability Pitfalls in Language Learning Systems
Ben Knight, Wm. Matthew Kennedy, Danielle Carvalho +2
AI-powered language learning tools increasingly provide instant, personalised feedback to millions of learners worldwide. However, this feedback can fail in ways that are difficult…
Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench
James Edgell, Wm. Matthew Kennedy, Isaac Pattis +3
The rapid adoption of large language models in AI-powered language education has created an urgent need for evaluations that assess pedagogical effectiveness, particularly in langu…