8 papers
PHF: Privileged Hidden Flow for On-Policy Self-Distillation
Yuhan Li, Mingxu Zhang, Dazhong Shen +1
On-policy self-distillation (OPSD) trains a reasoning model on rollouts sampled from its own policy by matching a privileged teacher that also sees verified reference solutions. Ex…
IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage
Yuhan Li, Mingxu Zhang, Dazhong Shen +1
Reinforcement learning with verifiable rewards (RLVR) has become a key technique for en- hancing LLM reasoning, yet its data ineffi- ciency remains a major bottleneck. Existing met…
SAE-FD: Sparse Autoencoder Feature Distillation for Continual Learning of Large Language Models
Mingxu Zhang, Yuhan Li, Lujundong Li +3
Continual learning enables large language models to adapt to evolving tasks without retraining from scratch, yet catastrophic forgetting remains a central obstacle. Among continual…
SLIM: Sparse Latent Steering for Interpretable and Property-Directed LLM-Based Molecular Editing
Mingxu Zhang, Yuhan Li, Lujundong Li +3
Large language models possess strong chemical reasoning capabilities, making them effective molecular editors. However, property-relevant information is implicitly entangled across…
ChemATP: A Training-Free Chemical Reasoning Framework for Large Language Models
Mingxu Zhang, Dazhong Shen, Qi Zhang +1
Large Language Models (LLMs) exhibit strong general reasoning but struggle in molecular science due to the lack of explicit chemical priors in standard string representations. Curr…
AtomDisc: An Atom-level Tokenizer that Boosts Molecular LLMs and Reveals Structure--Property Associations
Mingxu Zhang, Dazhong Shen, Ying Sun
Advances in large language models (LLMs) are accelerating discovery in molecular science. However, adapting molecular information to the serialized, token-based processing of LLMs…