9 papers
From Shallow to Deep: Pinning Semantic Intent via Causal GRPO
Shuyi Zhou, Zeen Song, Wenwen Qiang +4
Large Language Models remain vulnerable to adversarial prefix attacks (e.g., ``Sure, here is'') despite robust standard safety. We diagnose this vulnerability as Shallow Safety Ali…
Adaptive Uncertainty-Aware Tree Search for Robust Reasoning
Zeen Song, Zihao Ma, Wenwen Qiang +2
Inference-time reasoning scaling has significantly advanced the capabilities of Large Language Models (LLMs) in complex problem-solving. A prevalent approach involves external sear…
Causal Front-Door Adjustment for Robust Jailbreak Attacks on LLMs
Yao Zhou, Zeen Song, Wenwen Qiang +4
Safety alignment mechanisms in Large Language Models (LLMs) often operate as latent internal states, obscuring the model's inherent capabilities. Building on this observation, we m…
Group Causal Policy Optimization for Post-Training Large Language Models
Ziyin Gu, Jingyao Wang, Ran Zuo +4
Recent advances in large language models (LLMs) have broadened their applicability across diverse tasks, yet specialized domains still require targeted post training. Among existin…
Causal Reward Adjustment: Mitigating Reward Hacking in External Reasoning via Backdoor Correction
Ruike Song, Zeen Song, Huijie Guo +1
External reasoning systems combine language models with process reward models (PRMs) to select high-quality reasoning paths for complex tasks such as mathematical problem solving.…
Reward Model Generalization for Compute-Aware Test-Time Reasoning
Zeen Song, Wenwen Qiang, Siyu Zhao +2
External test-time reasoning enhances large language models (LLMs) by decoupling generation and selection. At inference time, the model generates multiple reasoning paths, and an a…