collaborators
Showing cs.CLShow all

7 papers · 1 filter

cs.CL2026

CurveShift: Is Agent Progress Scalar? Separating Level from Shape

Hanwen Xing, Pengyun Wang, BingXu Meng +8

Progress in large language models is often summarized using a single scalar measure, such as a time horizon, a latent ability estimate, or an aggregate benchmark score. These summa…

cs.CL2026

Co-Evolving Skill Generation and Policy Optimization

Zhiwei Zhang, Yudi Lin, Nikki Lijing Kuang +4

Skill-augmented reinforcement learning improves language agents by storing reusable procedural knowledge acquired from past experience. Existing methods typically use strong langua…

cs.CL2025

ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models

Xiaomin Li, Xupeng Chen, Jingxuan Fan +2

The safety alignment of large language models (LLMs) often relies on reinforcement learning from human feedback (RLHF), which requires human annotations to construct preference dat…

cs.CL2025

Selection of LLM Fine-Tuning Data based on Orthogonal Rules

Xiaomin Li, Mingye Gao, Zhiwei Zhang +2

High-quality training data is critical to the performance of large language models (LLMs). Recent work has explored using LLMs to rate and select data based on a small set of human…

cs.CL2025

MedGUIDE: Benchmarking Clinical Decision-Making in Large Language Models

Xiaomin Li, Mingye Gao, Yuexing Hao +4

Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making.…

cs.CL2025

Catastrophic Failure of LLM Unlearning via Quantization

Zhiwei Zhang, Fali Wang, Xiaomin Li +6

Large language models (LLMs) have shown remarkable proficiency in generating text, benefiting from extensive training on vast textual corpora. However, LLMs may also acquire unwant…