Showing cs.LGShow all
3 papers · 1 filter
cs.LG2026
Continual Learning in Transition
Zhiyan Hou, Dan Zhang, Tao Feng +11
Classical continual learning (CL) has primarily focused on enabling models to update and retain knowledge through parameter-centric mechanisms, e.g., training strategies, architect…
cs.LG2026
CompassDPO: Dynamics-Controlled Direct Preference Optimization for Robust Safety Alignment
Jilong Liu, Yonghui Yang, Pengyang Shao +5
Direct Preference Optimization (DPO) has become a standard framework for safety alignment, but its reliance on pairwise preference updates makes training sensitive to imperfect sup…
cs.LG2025
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
Junfeng Fang, Zijun Yao, Ruipeng Wang +3
The development of large language models (LLMs) has entered in a experience-driven era, flagged by the emergence of environment feedback-driven learning via reinforcement learning…