12 papers
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning
Jingyao Wang, Peizheng Guo, Wenwen Qiang +4
Large language models (LLMs) excel at complex tasks with advances in reasoning capabilities. However, existing reward mechanisms remain tightly coupled to final correctness and pay…
On the Plasticity and Stability for Post-Training Large Language Models
Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4
Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…
Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
Yao Zhou, Zeen Song, Wenwen Qiang +4
Safety alignment mechanisms in Large Language Models (LLMs) often operate as latent internal states, obscuring the model's inherent capabilities. Building on this observation, we m…
Self-Supervised Video Representation Learning in a Heuristic Decoupled Perspective
Zeen Song, Wenwen Qiang, Changwen Zheng +2
Video contrastive learning (V-CL) has emerged as a popular framework for unsupervised video representation learning, demonstrating strong results in tasks such as action classifica…
Understanding Token-level Topological Structures in Transformer-based Time Series Forecasting
Jianqi Zhang, Wenwen Qiang, Jingyao Wang +3
Transformer-based methods have achieved state-of-the-art performance in time series forecasting (TSF) by capturing positional and semantic topological relationships among input tok…
Learning to Think: Information-Theoretic Reinforcement Fine-Tuning for LLMs
Jingyao Wang, Wenwen Qiang, Zeen Song +2
Large language models (LLMs) excel at complex tasks thanks to advances in their reasoning abilities. However, existing methods overlook the trade-off between reasoning effectivenes…