activity
20242026
collaborators

8 papers

cs.CL2026

Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation

Keqin Peng, Chen Li, Yuanxin Ouyang +2

On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxicall…

cs.LG2026

VAE-Inf: A statistically interpretable generative paradigm for imbalanced classification

Hongfei Wu, Ruijian Han, Yancheng Yuan

Imbalanced classification remains a pervasive challenge in machine learning, particularly when minority samples are too scarce to provide a robust discriminative boundary. In such…

cs.CL2026

Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning

Keqin Peng, Yuanxin Ouyang, Xuebo Liu +4

Reinforcement Learning with Verifiable Rewards (RLVR) can elicit strong multi-step reasoning, yet it often encourages overly verbose traces. Moreover, naive length penalties in gro…

cs.LG2025

Accelerating RLHF Training with Reward Variance Increase

Zonglin Yang, Zhexuan Gu, Houduo Qi +1

Reinforcement learning from human feedback (RLHF) is an essential technique for ensuring that large language models (LLMs) are aligned with human values and preferences during the…

cs.CL2025

Enhancing Input-Label Mapping in In-Context Learning with Contrastive Decoding

Keqin Peng, Liang Ding, Yuanxin Ouyang +3

Large language models (LLMs) excel at a range of tasks through in-context learning (ICL), where only a few task examples guide their predictions. However, prior research highlights…

math.OC2025

PyClustrPath: An efficient Python package for generating clustering paths with GPU acceleration

Hongfei Wu, Yancheng Yuan

Convex clustering is a popular clustering model without requiring the number of clusters as prior knowledge. It can generate a clustering path by continuously solving the model wit…