collaborators

5 papers

cs.CL2026

Large Language Models Could Be Rote Learners

Yuyang Xu, Renjun Hu, Haochao Ying +3

Benchmark-based evaluation, e.g., multiple-choice questions (MCQs) and open-ended questions (OEQs), is widely used for evaluating Large Language Models (LLMs), yet their reliabilit…

cs.AI2025

Versatile and Risk-Sensitive Cardiac Diagnosis via Graph-Based ECG Signal Representation

Yue Wang, Yuyang Xu, Renjun Hu +7

Despite the rapid advancements of electrocardiogram (ECG) signal diagnosis and analysis methods through deep learning, two major hurdles still limit their clinical adoption: the la…

cs.LG2025

SSPO: Self-traced Step-wise Preference Optimization for Process Supervision and Reasoning Compression

Yuyang Xu, Yi Cheng, Haochao Ying +5

Test-time scaling has proven effective in further enhancing the performance of pretrained Large Language Models (LLMs). However, mainstream post-training methods (i.e., reinforceme…

cs.CL2025

LLMs Can Simulate Standardized Patients via Agent Coevolution

Zhuoyun Du, Lujie Zheng, Renjun Hu +7

Training medical personnel using standardized patients (SPs) remains a complex challenge, requiring extensive domain expertise and role-specific practice. Previous research on Larg…

cs.CL2025

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons

Renjun Hu, Yi Cheng, Libin Meng +4

The rapid advancement of large language models (LLMs) has opened new possibilities for their adoption as evaluative judges. This paper introduces Themis, a fine-tuned LLM judge tha…