activity
20242026
collaborators

9 papers

cs.CL2026

Learnable Assessment Skills for LLM-based Automated Scoring: Rubric Construction via Iterative Optimization

Yun Wang, Xin Xia, Xuansheng Wu +2

LLM-based automated scoring approaches near-human performance, but scaling to new tasks remains bottlenecked by the per-item human configuration of upstream stages such as rubric c…

cs.AI2026

Multi-Agent Causal Discovery Using Large Language Models

Hao Duong Le, Xin Xia, Haijie Xu +1

Causal discovery aims to identify causal relationships between variables and is a fundamental problem across the sciences. Traditional statistical causal discovery (SCD) methods re…

cs.LG2026

Heterogeneous Agent Collaborative Reinforcement Learning

Zhixia Zhang, Zixuan Huang, Gongxun Li +10

We introduce Heterogeneous Agent Collaborative Reinforcement Learning (HACRL), a new Reinforcement Learning from Verifiable Reward (RLVR) problem that addresses the inefficiencies…

cs.CL2026

Using Learning Progressions to Guide AI Feedback for Science Learning

Xin Xia, Nejla Yuruk, Yun Wang +1

Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubrics authored by domain experts. W…

cs.SE2025

AXIOM: Benchmarking LLM-as-a-Judge for Code via Rule-Based Perturbation and Multisource Quality Calibration

Ruiqi Wang, Xinchen Wang, Cuiyun Gao +3

Large language models (LLMs) have been increasingly deployed in real-world software engineering, fostering the development of code evaluation metrics to study the quality of LLM-ge…

cs.SE2025

An Empirical Study of Knowledge Distillation for Code Understanding Tasks

Ruiqi Wang, Zezhou Yang, Cuiyun Gao +2

Pre-trained language models (PLMs) have emerged as powerful tools for code understanding. However, deploying these PLMs in large-scale applications faces practical challenges due t…