5 citations · 13 across the 16 of their papers we have counts for
10 papers · 1 filter
Can GPT-4 do L2 analytic assessment?
Stefano Bannò, Hari Krishna Vydana, Kate M. Knill +1
Automated essay scoring (AES) to evaluate second language (L2) proficiency has been a firmly established technology used in educational contexts for decades. Although holistic scor…
Question Difficulty Ranking for Multiple-Choice Reading Comprehension
Vatsal Raina, Mark Gales
Multiple-choice (MC) tests are an efficient method to assess English learners. It is useful for test creators to rank candidate MC questions by difficulty during exam curation. Typ…
WaterJudge: Quality-Detection Trade-off when Watermarking Large Language Models
Piotr Molenda, Adian Liusie, Mark J. F. Gales
Watermarking generative-AI systems, such as LLMs, has gained considerable interest, driven by their enhanced capabilities across a wide range of tasks. Although current approaches…
Teacher-Student Training for Debiasing: General Permutation Debiasing for Large Language Models
Adian Liusie, Yassir Fathullah, Mark J. F. Gales
Large Language Models (LLMs) have demonstrated impressive zero-shot capabilities and versatility in NLP tasks, however they sometimes fail to maintain crucial invariances for speci…
An Information-Theoretic Approach to Analyze NLP Classification Tasks
Luran Wang, Mark Gales, Vatsal Raina
Understanding the importance of the inputs on the output is useful across many tasks. This work provides an information-theoretic framework to analyse the influence of inputs for t…
Assessing Distractors in Multiple-Choice Tests
Vatsal Raina, Adian Liusie, Mark Gales
Multiple-choice tests are a common approach for assessing candidates' comprehension skills. Standard multiple-choice reading comprehension exams require candidates to select the co…