6 papers · 1 filter
Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation
Kaiqi Yang, Tai-Quan Peng, Sanguk Lee +1
LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frame…
Optimizing In-Context Demonstrations for LLM-based Automated Grading
Yucheng Chu, Hang Li, Kaiqi Yang +4
Automated assessment of open-ended student responses is a critical capability for scaling personalized feedback in education. While large language models (LLMs) have shown promise…
Confusion-Aware Rubric Optimization for LLM-based Automated Grading
Yucheng Chu, Hang Li, Kaiqi Yang +4
Accurate and unambiguous guidelines are critical for large language model (LLM) based graders, yet manually crafting these prompts is often sub-optimal as LLMs can misinterpret exp…
How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment
Hang Li, Kaiqi Yang, Xianxuan Long +9
The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptabili…
A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization
Yucheng Chu, Hang Li, Kaiqi Yang +4
Open-ended short-answer questions (SAGs) have been widely recognized as a powerful tool for providing deeper insights into learners' responses in the context of learning analytics…
Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations
Sanguk Lee, Kai-Qi Yang, Tai-Quan Peng +2
Large language models (LLMs) are employed to simulate human-like responses in social surveys, yet it remains unclear if they develop biases like social desirability response (SDR)…