6 papers
Continual Learning in Transition
Zhiyan Hou, Dan Zhang, Tao Feng +11
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architect…
ResMerge: Residual-based Spectral Merging of Large Language Models
Yandu Sun, Zhiyan Hou, Haokai Ma +6
Model merging offers a training-free way to combine multiple post-trained expert models, but merging experts obtained through reinforcement learning (RL) remains challenging. Exist…
TRACE: Trajectory Risk-Aware Compression for Long-Horizon Agent Safety
Zhepei Hong, Lin Wang, Liting Li +5
Long-horizon LLM agents produce safety evidence across long trajectories, where sparse, delayed, and compositional risk signals often escape local moderation. Existing turn-level o…
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
Jilong Liu, Yonghui Yang, Pengyang Shao +5
Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…
DualEdit: Mitigating Safety Fallback in LLM Backdoor Editing via Affirmation-Refusal Regulation
Houcheng Jiang, Zetong Zhao, Junfeng Fang +5
Safety-aligned large language models (LLMs) remain vulnerable to backdoor attacks. Recent model editing-based approaches enable efficient backdoor injection by directly modifying a…
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
Junfeng Fang, Zijun Yao, Ruipeng Wang +3
The development of large language models (LLMs) has entered in a experience-driven era, flagged by the emergence of environment feedback-driven learning via reinforcement learning…