From the 1 of 4 linked papers with an AI index.
4 papers
How Well Does AI-Generated Feedback Work? Intrinsic and Extrinsic Evaluation across more than 20,000 EFL Essay Drafts
Steven Coyne, Diana Galvan-Sosa, Ryan Spring +4
The paper investigates AI-generated written corrective feedback for English‑as‑a‑Foreign‑Language essays, comparing teacher (intrinsic) ratings with student (extrinsic) responses a…
Expert Evaluation of LLM's Open-Ended Legal Reasoning on the Japanese Bar Exam Writing Task
Jungmin Choi, Keisuke Sakaguchi, Hiroaki Yamada
Large language models (LLMs) have shown strong performance on legal benchmarks, including multiple-choice components of bar exams. However, their capacity for generating open-ended…
Annotating Errors in English Learners' Written Language Production: Advancing Automated Written Feedback Systems
Steven Coyne, Diana Galvan-Sosa, Ryan Spring +4
Recent advances in natural language processing (NLP) have contributed to the development of automated writing evaluation (AWE) systems that can correct grammatical errors. However,…
Rubrik's Cube: Testing a New Rubric for Evaluating Explanations on the CUBE dataset
Diana Galvan-Sosa, Gabrielle Gaudeau, Pride Kavumba +5
The performance and usability of Large-Language Models (LLMs) are driving their use in explanation generation tasks. However, despite their widespread adoption, LLM explanations ha…