AI evaluation 1benchmarking 1education assessment 1large language models 1second language learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CY2026
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight +3
The paper presents L2-Bench, an open‑source benchmark of over 1,000 task‑response pairs designed to evaluate large language models on competencies relevant to second language (L2)…
cs.CY2026
Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench
James Edgell, Wm. Matthew Kennedy, Isaac Pattis +3
The rapid adoption of large language models in AI-powered language education has created an urgent need for evaluations that assess pedagogical effectiveness, particularly in langu…