18 papers
Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
Luyang Fang, Yingchuan Zhang, Jongchan Park +3
Student-generated drawings are widely used in science education to assess learners' conceptual understanding in modeling-based tasks aligned with the Next Generation Science Standa…
Trust the Right Teacher: Quality-Aware Self-Distillation for GUI Grounding
Jingyuan Huang, Zuming Huang, Yucheng Shi +4
Graphical user interface (GUI) grounding requires vision-language models (VLMs) to identify small target elements in high-resolution screenshots and predict precise screen coordina…
Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization
Yun Wang, Xin Xia, Xuansheng Wu +2
LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric c…
Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment
Yizhu Gao, Zhongzhou Chen, Min Li +1
Generative Artificial Intelligence (GenAI) presents a governance challenge for STEM assessment. Unrestricted GenAI access enables task outsourcing that undermines the validity of t…
Palette: A Modular, Controllable, and Efficient Framework for On-demand Authorized Safety Alignment Relaxation in LLMs
Qitao Tan, Xiaoying Song, Arman Akbari +7
Current safety alignment of foundation models largely follows a \emph{one-size-fits-all} paradigm, applying the same refusal policy across users and contexts. As a result, models m…
BRIDGE the Gap: Mitigating Bias Amplification in Automated Scoring of English Language Learners via Inter-group Data Augmentation
Yun Wang, Xuansheng Wu, Jingyuan Huang +3
In the field of educational assessment, automated scoring systems increasingly rely on deep learning and large language models (LLMs). However, these systems face significant risks…