8 papers
Diagnosing and Mitigating Thinking Collapse in On-Policy Self-Distillation
Keqin Peng, Chen Li, Yuanxin Ouyang +2
On-Policy Self-Distillation (OPSD) has emerged as a crucial paradigm for enhancing and aligning Large Language Models (LLMs). However, in complex reasoning tasks, OPSD paradoxicall…
VAE-Inf: A statistically interpretable generative paradigm for imbalanced classification
Hongfei Wu, Ruijian Han, Yancheng Yuan
Imbalanced classification remains a pervasive challenge in machine learning, particularly when minority samples are too scarce to provide a robust discriminative boundary. In such…
Think Dense, Not Long: Dynamic Decoupled Conditional Advantage for Efficient Reasoning
Keqin Peng, Yuanxin Ouyang, Xuebo Liu +4
Reinforcement Learning with Verifiable Rewards (RLVR) can elicit strong multi-step reasoning, yet it often encourages overly verbose traces. Moreover, naive length penalties in gro…
Accelerating RLHF Training with Reward Variance Increase
Zonglin Yang, Zhexuan Gu, Houduo Qi +1
Reinforcement learning from human feedback (RLHF) is an essential technique for ensuring that large language models (LLMs) are aligned with human values and preferences during the…
Enhancing Input-Label Mapping in In-Context Learning with Contrastive Decoding
Keqin Peng, Liang Ding, Yuanxin Ouyang +3
Large language models (LLMs) excel at a range of tasks through in-context learning (ICL), where only a few task examples guide their predictions. However, prior research highlights…
PyClustrPath: An efficient Python package for generating clustering paths with GPU acceleration
Hongfei Wu, Yancheng Yuan
Convex clustering is a popular clustering model without requiring the number of clusters as prior knowledge. It can generate a clustering path by continuously solving the model wit…