AI evaluation 1benchmarking 1education assessment 1large language models 1second language learning 1
From the 1 of 2 linked papers with an AI index.
2 papers
cs.CY2026
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight +3
The paper presents L2-Bench, an open‑source benchmark of over 1,000 task‑response pairs designed to evaluate large language models on competencies relevant to second language (L2)…
cs.AI2026
Every Eval Ever: A Unifying Schema and Community Repository for AI Evaluation Results
Jan Batzner, Sree Harsha Nelaturu, Damian Stachura +45
AI evaluations are widely used for testing and understanding progress. However, the diverse evaluators bring with them inconsistencies that challenge analysis and comparison. First…