collaborators

5 papers

cs.CL2026

Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs

Xinyi Wang, Hong Jiao, Ming Li +4

The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LL…

cs.CL2025

Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models

Qingshu Xu, Hong Jiao, Tianyi Zhou +4

Accurate alignment of items to content standards is critical for valid score interpretation in large-scale assessments. This study evaluates three automated paradigms for aligning…

cs.CL2025

Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests

Yanbin Fu, Hong Jiao, Tianyi Zhou +5

Aligning test items to content standards is a critical step in test development to collect validity evidence based on content. Item alignment has typically been conducted by human…

cs.CL2025

Text-Based Approaches to Item Difficulty Modeling in Large-Scale Assessments: A Systematic Review

Sydney Peters, Nan Zhang, Hong Jiao +3

Item difficulty plays a crucial role in test performance, interpretability of scores, and equity for all test-takers, especially in large-scale assessments. Traditional approaches…

cs.AI2025

Understanding the Thinking Process of Reasoning Models: A Perspective from Schoenfeld's Episode Theory

Ming Li, Nan Zhang, Chenrui Fan +6

While Large Reasoning Models (LRMs) generate extensive chain-of-thought reasoning, we lack a principled framework for understanding how these thoughts are structured. In this paper…