3 papers
cs.LG2026
When Autoregressive Consistency Hurts Safety Alignment
Bochen Lyu, Yiyang Jia, Xiaohao Cai +1
Safety alignment in large language models (LLMs) is fragile in part because it is often shallow: fine-tuning mainly reshapes the model's behavior near the first few output tokens.…
cs.LG2026
Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently
Bochen Lyu, Yiyang Jia, Xiaohao Cai +1
Transformers can acquire Chain-of-Thought (CoT) capabilities to solve reasoning tasks via fine-tuning. Reinforcement learning (RL) and supervised fine-tuning (SFT) are two primary…
cs.LG2025
Heavy-Ball Momentum Method in Continuous Time and Discretization Error Analysis
Bochen Lyu, Xiaojing Zhang, Fangyi Zheng +3
This paper establishes a continuous time approximation, a piece-wise continuous differential equation, for the discrete Heavy-Ball (HB) momentum method with explicit discretization…