activity
20242026
collaborators

9 papers

cs.CL2026

EDIT: Evidence-Diagnosed Intervention Training for Rule-Faithful LLM Grading

Zhihao Wu, Linhai Zhang, Taiyi Wang +4

Reliable rubric grading requires more than accurate score prediction. Each judgement must be grounded in the mark scheme and evidence from the student answer. Existing credit-assig…

cs.CL2026

Staying with the Uncertainty: Uncertainty-Scaffolding Strategies for Artificial Moral Advisors in LLM-to-LLM Simulated Conversations

Salvatore Greco, Hainiu Xu, Jacopo Domenicucci +2

LLMs are increasingly deployed as Artificial Moral Advisors (AMA) in a variety of contexts: what kind of conversational patterns should they display? In this paper, we study how AM…

cs.CL2026

SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

Zhenyi Shen, Junru Lu, Lin Gui +4

Sparse attention reduces the quadratic complexity of full self-attention but faces two challenges: (1) an attention gap, where applying sparse attention to full-attention-trained m…

cs.CL2025

Two Heads Are Better Than One: Dual-Model Verbal Reflection at Inference-Time

Jiazheng Li, Yuxiang Zhou, Junru Lu +4

Although preference optimization methods have improved reasoning performance in Large Language Models (LLMs), they often lack transparency regarding why one reasoning outcome is pr…

cs.CL2025

AERA Chat: An Interactive Platform for Automated Explainable Student Answer Assessment

Jiazheng Li, Artem Bobrov, Runcong Zhao +2

Explainability in automated student answer scoring systems is critical for building trust and enhancing usability among educators. Yet, generating high-quality assessment rationale…

cs.CL2025

RoleMRC: A Fine-Grained Composite Benchmark for Role-Playing and Instruction-Following

Junru Lu, Jiazheng Li, Guodong Shen +5

Role-playing is important for Large Language Models (LLMs) to follow diverse instructions while maintaining role identity and the role's pre-defined ability limits. Existing role-p…