1 citations · 1 across the 6 of their papers we have counts for
Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Subliminal Learning as Trait-Direction Drift: A Mechanism and Targeted Control under SFT Distillation
Zhixuan Liu, Zhichen Dong, Yuyu Fan +2
Beyond intended capabilities, model distillation can transfer hidden traits from a teacher. A teacher biased by a system prompt can generate semantically clean training data, such…
cs.LG2026
Scaling Model-Generated Distillation Data Can Make Latent Teacher Traits More Recoverable
Zhichen Dong, Zhixuan Liu, Yuyu Fan +3
Scaling model-generated data is usually viewed as improving distillation: more examples should increase coverage, reduce noise, and produce stronger students. We show a second effe…
cs.LG2026
How Does Reasoning Flow? Tracing Attention-Induced Information Flow for Targeted RL in LLMs
Zhichen Dong, Yang Li, Yuhan Sun +9
Token-level credit assignment remains a key obstacle for reinforcement learning (RL) in large language models (LLMs), where RL recipes typically treat all tokens equally, failing t…