2 papers
cs.LG2026
When Autoregressive Consistency Hurts Safety Alignment
Bochen Lyu, Yiyang Jia, Xiaohao Cai +1
Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens.…
cs.LG2025
Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis
Bochen Lyu, Xiaojing Zhang, Fangyi Zheng +3
This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization…