30 papers
From Generic Correlation to Input-Specific Credit in On-Policy Self Distillation
Guobin Shen, Lei Huang, Xiang Cheng +4
On-policy self-distillation has emerged as a promising paradigm for post-training language models, in which the model conditions on environment feedback to serve as its own teacher…
Anti-Self-Distillation for Reasoning RL via Pointwise Mutual Information
Guobin Shen, Xiang Cheng, Chenxiao Zhao +4
On-policy self-distillation, where a student is pulled toward a copy of itself conditioned on privileged context (e.g., a verified solution or feedback), offers a promising directi…
Efficient LLM Safety Evaluation through Multi-Agent Debate
Dachuan Lin, Guobin Shen, Zihao Yang +3
Safety evaluation of large language models (LLMs) increasingly relies on LLM-as-a-judge pipelines, but strong judges can still be expensive to use at scale. We study whether struct…
Towards Reliable Evaluation of Adversarial Robustness for Spiking Neural Networks
Jihang Wang, Dongcheng Zhao, Ruolin Chen +2
Spiking Neural Networks (SNNs) utilize spike-based activations to mimic the brain's energy-efficient information processing. However, the binary and discontinuous nature of spike a…
Light Alignment Improves LLM Safety via Model Self-Reflection with a Single Neuron
Sicheng Shen, Mingyang Lv, Han Shen +7
The safety of large language models (LLMs) has increasingly emerged as a fundamental aspect of their development. Existing safety alignment for LLMs is predominantly achieved throu…
Multi-Level Safety Continual Projection for Fine-Tuned Large Language Models without Retraining
Bing Han, Feifei Zhao, Dongcheng Zhao +4
While fine-tuning services drive the rapid expansion of task capabilities in large language models (LLMs), they are often accompanied by the degradation and reorganization of safet…