7 papers
Constitutional On-Policy Safe Distillation
Ming Wen, Yuxuan Liu, Kun Yang +9
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervis…
Unraveling the Hidden Dynamical Structure in Recurrent Neural Policies
Jin Li, Yue Wu, Mengsha Huang +3
Recurrent neural policies are widely used in partially observable control and meta-RL tasks. Their abilities to maintain internal memory and adapt quickly to unseen scenarios have…
Orthogonal Concept Erasure for Diffusion Models
Yuhao Sun, Lingyun Yu, Haoxiang Xu +3
Concept erasure has emerged as a promising approach to mitigate undesired or unsafe content in diffusion models, yet existing methods still face significant limitations. While trai…
Be Careful When Fine-tuning On Open-Source LLMs: Your Fine-tuning Data Could Be Secretly Stolen!
Zhexin Zhang, Yuhao Sun, Junxiao Yang +4
Fine-tuning on open-source Large Language Models (LLMs) with proprietary data is now a standard practice for downstream developers to obtain task-specific LLMs. Surprisingly, we re…
Multiplicative Orthogonal Sequential Editing for Language Models
Hao-Xiang Xu, Jun-Yu Ma, Ziqi Peng +3
Knowledge editing aims to efficiently modify the internal knowledge of large language models (LLMs) without compromising their other capabilities. The prevailing editing paradigm,…
UpSafeC: Upcycling for Controllable Safety in Large Language Models
Yuhao Sun, Zhuoer Xu, Shiwen Cui +4
Large Language Models (LLMs) have achieved remarkable progress across a wide range of tasks, but remain vulnerable to safety risks such as harmful content generation and jailbreak…