collaborators

13 papers

cs.CL2026

ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models

Yilin Jiang, Xiaorong Zhu, Fei Tan +9

Large language models are increasingly deployed in education as tutors, teaching assistants, and content generators. These roles place demands that ordinary question answering does…

cs.AI2026

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents

Haojie Hao, Longkun Hao, Yihang Lou +8

Reinforcement Learning (RL) has become a promising approach for improving GUI Agents in long-horizon, stochastic digital environments, but trajectory-level success feedback is too…

cs.LG2026

Speculative Rollback Correction for Quality-Diverse Web Agent Imitation

Longkun Hao, Hongyu Lin, Hao Li +11

Training interactive web agents through imitation learning from expert trajectories has emerged as a highly effective approach. However, determining the optimal timing for expert i…

cs.LG2026

Elmes*: Automated Construction of Fine-Grained Evaluation Rubrics for Large Language Models in Long-Tail Educational Scenarios

Tao Liu, Ye Lu, Ruohua Zhang +4

Evaluating large language models (LLMs) for education requires measuring how models teach, not only what they know. Existing benchmarks emphasize domain-general correctness or depe…

cs.AI2026

Automating Skill Acquisition through Large-Scale Mining of Open-Source Agentic Repositories: A Framework for Multi-Agent Procedural Knowledge Extraction

Shuzhen Bi, Mengsong Wu, Hao Hao +5

The transition from monolithic large language models (LLMs) to modular, skill-equipped agents represents a fundamental architectural shift in artificial intelligence deployment. Wh…

cs.AI2026

Scaling Laws for Educational AI Agents

Mengsong Wu, Hao Hao, Shuzhen Bi +5

While scaling laws for Large Language Models (LLMs) have been extensively studied along dimensions of model parameters, training data, and compute, the scaling behavior of LLM-base…