8 papers
Representing Visual Evidence for Item Difficulty Prediction: Visual Textualization and Image-Native Modeling
Han Chen, Ming Li, Hong Jiao +1
Predicting item difficulty from content can provide an initial estimate for newly developed questions before sufficient student responses are available. Existing approaches typical…
LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment
Han Chen, Ming Li, Chenguang Wang +4
Existing work on LLM-based educational assessment has focused largely on item difficulty, but difficulty alone does not indicate whether an item meaningfully distinguishes higher-…
Can LLMs Really Understand Item Difficulty Levels? Implications for Automated Item Generation Using LLMs
Xinyi Wang, Hong Jiao, Ming Li +4
The estimation of item difficulty plays a key role in both formative assessment and large-scale high-stakes summative assessments. This study explores how large language models (LL…
Can LLMs Estimate Student Struggles? Human-AI Difficulty Alignment with Proficiency Simulation for Item Difficulty Prediction
Ming Li, Han Chen, Yunze Xiao +3
Accurate estimation of item (question or task) difficulty is critical for educational assessment but suffers from the cold start problem. While Large Language Models demonstrate su…
Automated Alignment of Math Items to Content Standards in Large-Scale Assessments Using Language Models
Qingshu Xu, Hong Jiao, Tianyi Zhou +4
Accurate alignment of items to content standards is critical for valid score interpretation in large-scale assessments. This study evaluates three automated paradigms for aligning…
Text-Based Approaches to Item Alignment to Content Standards in Large-Scale Reading & Writing Tests
Yanbin Fu, Hong Jiao, Tianyi Zhou +5
Aligning test items to content standards is a critical step in test development to collect validity evidence based on content. Item alignment has typically been conducted by human…