3 papers
cs.CL2026
Learning What to Learn: Stage-Specific Data Sets for SFT-then-RL in Small Language Model Reasoning
Chongyang He, Rui Zhang, Zixuan Wang +1
Post-training Small Language Models (SLMs) for reasoning typically follows an SFT-then-RL pipeline, yet existing work rarely considers what data should be learned at each stage. We…
cs.SD2026
CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding
Eugene Kwek, Feng Liu, Rui Zhang +1
Neural audio codecs are a key component of speech processing pipelines, compressing audio into discrete tokens for downstream modeling. However, existing codecs struggle to balance…
cs.LG2026
Representation Collapse in Sequential Post-Training of Large Language Models
Yichen Liu, Mingyu Chen, Hao Wang +7
Large language models are now adapted through chains of post-training stages rather than through a single instruction-tuning pass. This paper studies whether such sequential post-t…