3 papers
cs.CY2026
L2-Bench: An Evaluation Benchmark for Measuring LLM Capabilities in Second Language Education
James Edgell, Wm. Matthew Kennedy, Ben Knight +3
Despite rapid AI adoption in education, rigorous evaluation of AI-powered educational (AIED) systems remains critically underdeveloped, particularly in second language (L2) educati…
cs.LG2026
Choosing the Right Regularizer for Applied ML: Simulation Benchmarks of Popular Scikit-learn Regularization Frameworks
Benjamin S. Knight, Ahsaas Bajaj
This study surveys the historical development of regularization, tracing its evolution from stepwise regression in the 1960s to recent advancements in formal error control, structu…
cs.CY2026
Towards an Evaluation Methodology for AI in Second Language Education: Lessons Learned from Developing L2-Bench
James Edgell, Wm. Matthew Kennedy, Isaac Pattis +3
The rapid adoption of large language models in AI-powered language education has created an urgent need for evaluations that assess pedagogical effectiveness, particularly in langu…