3 papers
cs.LG2026
Adaptive Uncertainty-Aware Tree Search for Robust Reasoning
Zeen Song, Zihao Ma, Wenwen Qiang +2
Inference-time reasoning scaling has significantly advanced the capabilities of Large Language Models (LLMs) in complex problem-solving. A prevalent approach involves external sear…
cs.LG2026
On the Plasticity and Stability for Post-Training Large Language Models
Wenwen Qiang, Ziyin Gu, Jiahuan Zhou +4
Training stability remains a critical bottleneck for Group Relative Policy Optimization (GRPO), often manifesting as a trade-off between reasoning plasticity and general capability…
cs.CL2026
Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
Yao Zhou, Zeen Song, Wenwen Qiang +4
Safety alignment mechanisms in Large Language Models (LLMs) often operate as latent internal states, obscuring the model's inherent capabilities. Building on this observation, we m…