Showing cs.LGShow all
2 papers · 1 filter
cs.LG2026
Constitutional On-Policy Safe Distillation
Ming Wen, Yuxuan Liu, Kun Yang +9
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervis…
cs.LG2026
Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs
Ming Wen, Kun Yang, Xin Chen +4
Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently gen…