5 papers
Constitutional On-Policy Safe Distillation
Ming Wen, Yuxuan Liu, Kun Yang +9
On-policy self-distillation (OPSD) has emerged as an efficient post-training paradigm by using a teacher conditioned on privileged information to provide dense token-level supervis…
OOD-MMSafe: Advancing MLLM Safety from Harmful Intent to Hidden Consequences
Ming Wen, Kun Yang, Jingyu Zhang +4
While safety alignment for Multimodal Large Language Models (MLLMs) has gained significant attention, current paradigms primarily target malicious intent or situational violations.…
Pragma-VL: Towards a Pragmatic Arbitration of Safety and Helpfulness in MLLMs
Ming Wen, Kun Yang, Xin Chen +4
Multimodal Large Language Models (MLLMs) pose critical safety challenges, as they are susceptible not only to adversarial attacks such as jailbreaking but also to inadvertently gen…
Omni-Safety under Cross-Modality Conflict: Vulnerabilities, Dynamics Mechanisms and Efficient Alignment
Kun Wang, Zherui Li, Zhenhong Zhou +8
Omni-modal Large Language Models (OLLMs) greatly expand LLMs' multimodal capabilities but also introduce cross-modal safety risks. However, a systematic understanding of vulnerabil…
CSR-Bench: A Benchmark for Evaluating the Cross-modal Safety and Reliability of MLLMs
Yuxuan Liu, Yuntian Shi, Kun Wang +2
Multimodal large language models (MLLMs) enable interaction over both text and images, but their safety behavior can be driven by unimodal shortcuts instead of true joint intent un…