4 papers
On the Plasticity and Stability for Post-Training Large Language Models
Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4
Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…
Group Causal Policy Optimization for Post-Training Large Language Models
Ziyin Gu, Jingyao Wang, Ran Zuo +4
Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post training. Among existin…
On the Transferability and Discriminability of Repersentation Learning in Unsupervised Domain Adaptation
Wenwen Qiang, Ziyin Gu, Lingyu Si +4
In this paper, we addressed the limitation of relying solely on distribution alignment and source-domain empirical risk minimization in Unsupervised Domain Adaptation (UDA). Our in…
On the Generalization and Causal Explanation in Self-Supervised Learning
Wenwen Qiang, Zeen Song, Ziyin Gu +4
Self-supervised learning (SSL) methods learn from unlabeled data and achieve high generalization performance on downstream tasks. However, they may also suffer from overfitting to…