7 papers
Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
Zhixuan Liu, Zhichen Dong, Yuyu Fan +2
Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such…
Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable
Zhichen Dong, Zhixuan Liu, Yuyu Fan +3
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effe…
Decoupled Contrastive Decoding via Expert-Aligned Drafting
Zhixuan Liu, Zhichen Dong, Yuanfu Wang +1
Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment qu…
How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
Zhichen Dong, Yang Li, Yuhan Sun +9
Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing t…
SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law
Shanghai AI Lab, :, Yicheng Bao +115
We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…
Emergent Response Planning in LLMs
Zhichen Dong, Zhanhui Zhou, Zhixuan Liu +2
In this work, we argue that large language models (LLMs), though trained to predict only the next token, exhibit emergent planning behaviors: $\textbf{their hidden representations…