1 paper
Robinson Ferrer, Damla Turgut, Zhongzhou Chen +1
Large Language Models (LLMs) show promise for automated grading, but their outputs can be unreliable. Rather than improving grading accuracy directly, we address a complementary pr…