activity
20242026
collaborators

7 papers

cs.LG2026

Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation

Zhixuan Liu, Zhichen Dong, Yuyu Fan +2

Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such…

cs.LG2026

Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable

Zhichen Dong, Zhixuan Liu, Yuyu Fan +3

Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effe…

cs.CL2026

Decoupled Contrastive Decoding via Expert-Aligned Drafting

Zhixuan Liu, Zhichen Dong, Yuanfu Wang +1

Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment qu…

cs.LG2026

How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs

Zhichen Dong, Yang Li, Yuhan Sun +9

Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing t…

cs.AI2025

SafeWork-R1: Coevolving Safety and Intelligence under the AI-45 Law

Shanghai AI Lab, :, Yicheng Bao +115

We introduce SafeWork-R1, a cutting-edge multimodal reasoning model that demonstrates the coevolution of capabilities and safety. It is developed by our proposed SafeLadder framewo…

cs.CL2025

Emergent Response Planning in LLMs

Zhichen Dong, Zhanhui Zhou, Zhixuan Liu +2

In this work, we argue that large language models (LLMs), though trained to predict only the next token, exhibit emergent planning behaviors: $\textbf{their hidden representations…