7 papers
Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
Zhixuan Liu, Zhichen Dong, Yuyu Fan +2
Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such…
Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable
Zhichen Dong, Zhixuan Liu, Yuyu Fan +3
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effe…
Decoupled Contrastive Decoding via Expert-Aligned Drafting
Zhixuan Liu, Zhichen Dong, Yuanfu Wang +1
Contrastive Decoding (CD) improves generation quality, but its amateur-model pass makes decoding expensive. Accelerating CD with speculative decoding raises a proposal-alignment qu…
Native Reasoning Models: Training Language Models to Reason on Unverifiable Data
Yuanfu Wang, Zhixuan Liu, Xiangtian Li +2
The prevailing paradigm for training large reasoning models--combining Supervised Fine-Tuning (SFT) with Reinforcement Learning with Verifiable Rewards (RLVR)--is fundamentally con…
Emergent Response Planning in LLMs
Zhichen Dong, Zhanhui Zhou, Zhixuan Liu +2
In this work, we argue that large language models (LLMs), though trained to predict only the next token, exhibit emergent planning behaviors: $\textbf{their hidden representations…
Inference-Time Language Model Alignment via Integrated Value Guidance
Zhixuan Liu, Zhanhui Zhou, Yuanfu Wang +2
Large language models are typically fine-tuned to align with human preferences, but tuning large models is computationally intensive and complex. In this work, we introduce $\texti…