activity
20242026
collaborators
Showing cs.AIShow all

6 papers · 1 filter

cs.AI2026

Think-Before-Speak: From Internal Evaluation to Public Expression in Multi-Agent Social Simulation

Kaiqi Yang, Tai-Quan Peng, Sanguk Lee +1

LLM-based multi-agent simulation offers a promising way to study social interaction, deliberation, and collective opinion dynamics. However, many existing dialogue simulation frame…

cs.AI2026

Optimizing In-Context Demonstrations for LLM-based Automated Grading

Yucheng Chu, Hang Li, Kaiqi Yang +4

Automated assessment of open-ended student responses is a critical capability for scaling personalized feedback in education. While large language models (LLMs) have shown promise…

cs.AI2026

Confusion-Aware Rubric Optimization for LLM-based Automated Grading

Yucheng Chu, Hang Li, Kaiqi Yang +4

Accurate and unambiguous guidelines are critical for large language model (LLM) based graders, yet manually crafting these prompts is often sub-optimal as LLMs can misinterpret exp…

cs.AI2026

How Uncertain Is the Grade? A Benchmark of Uncertainty Metrics for LLM-Based Automatic Assessment

Hang Li, Kaiqi Yang, Xianxuan Long +9

The rapid rise of large language models (LLMs) is reshaping the landscape of automatic assessment in education. While these systems demonstrate substantial advantages in adaptabili…

cs.AI2024

Exploring Social Desirability Response Bias in Large Language Models: Evidence from GPT-4 Simulations

Sanguk Lee, Kai-Qi Yang, Tai-Quan Peng +2

Large language models (LLMs) are employed to simulate human-like responses in social surveys, yet it remains unclear if they develop biases like social desirability response (SDR)…

cs.AI2024

A LLM-Powered Automatic Grading Framework with Human-Level Guidelines Optimization

Yucheng Chu, Hang Li, Kaiqi Yang +4

Open-ended short-answer questions (SAGs) have been widely recognized as a powerful tool for providing deeper insights into learners' responses in the context of learning analytics…