1 paper
Jerin George Mathew, Sumayya Taher, Anindita Kundu +1
Large language models have recently been proposed as tools for automated essay scoring, but their agreement with human grading remains unclear. In this work, we evaluate how LLM-ge…