14 papers
On the Plasticity and Stability for Post-Training Large Language Models
Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4
Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…
Enhancing Large Language Models for Time-Series Forecasting via Vector-Injected In-Context Learning
Jianqi Zhang, Jingyao Wang, Wenwen Qiang +2
The World Wide Web needs reliable predictive capabilities to respond to changes in user behavior and usage patterns. Time series forecasting (TSF) is a key means to achieve this go…
Exploring Transferability of Self-Supervised Learning by Task Conflict Calibration
Huijie Guo, Jingyao Wang, Peizheng Guo +3
In this paper, we explore the transferability of SSL by addressing two central questions: (i) what is the representation transferability of SSL, and (ii) how can we effectively mod…
A Generalized Learning Framework for Self-Supervised Contrastive Learning
Lingyu Si, Jingyao Wang, Wenwen Qiang
Self-supervised contrastive learning (SSCL) has recently demonstrated superiority in multiple downstream tasks. In this paper, we generalize the standard SSCL methods to a Generali…
Group Causal Policy Optimization for Post-Training Large Language Models
Ziyin Gu, Jingyao Wang, Ran Zuo +4
Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post training. Among existin…
COPO: Causal-Oriented Policy Optimization for Hallucinations of MLLMs
Peizheng Guo, Jingyao Wang, Wenwen Qiang +3
Despite Multimodal Large Language Models (MLLMs) having shown impressive capabilities, they may suffer from hallucinations. Empirically, we find that MLLMs attend disproportionatel…