7 citations · 7 across the 4 of their papers we have counts for
Showing cs.CLShow all
3 papers · 1 filter
cs.CL2026
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
Yun Wang, Xin Xia, Xuansheng Wu +2
LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric c…
cs.CL2026
Using Learning Progressions to Guide AI Feedback for Science Learning
Xin Xia, Nejla Yuruk, Yun Wang +1
Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubrics authored by domain experts. W…
cs.CL2025
PHYBench: Holistic Evaluation of Physical Perception and Reasoning in Large Language Models
Shi Qiu, Shaoyang Guo, Zhuo-Yang Song +51
Current benchmarks for evaluating the reasoning capabilities of Large Language Models (LLMs) face significant limitations: task oversimplification, data contamination, and flawed e…