3 papers
cs.CL2025
How to Select Datapoints for Efficient Human Evaluation of NLG Models?
Vilém Zouhar, Peng Cui, Mrinmaya Sachan
Human evaluation is the gold standard for evaluating text generation models. However, it is expensive. In order to fit budgetary constraints, a random subset of the test data is of…
cs.CL2025
Grammar Control in Dialogue Response Generation for Language Learning Chatbots
Dominik Glandorf, Peng Cui, Detmar Meurers +1
Chatbots based on large language models offer cheap conversation practice opportunities for language learners. However, they are hard to control for linguistic forms that correspon…
cs.CL2025
Investigating the Zone of Proximal Development of Language Models for In-Context Learning
Peng Cui, Mrinmaya Sachan
In this paper, we introduce a learning analytics framework to analyze the in-context learning (ICL) behavior of large language models (LLMs) through the lens of the Zone of Proxima…