10 papers
ReFrame: Evidence-Guided Test-Time Safety Alignment in Multimodal Large Language Models
Wenzheng Jiang, Xuankun Rong, Yuanzhao Zhai +2
While multimodal large language models (MLLMs) extend model capabilities beyond text, they also make safety alignment increasingly challenging. Multimodal safety alignment methods…
Do LLMs Know Their Vulnerable Scenarios?
Ziheng Peng, Huiqi Deng, Haoran Jing +5
Safety-aligned large language models are trained to refuse harmful requests, yet embedding the same requests in particular scenarios can bypass their safeguards. Existing red-teami…
Behavioural Signatures of Risk-Sensitive Decision-Making in Large Language Models
Xuankun Rong, Wenke Huang, Bo Du +2
As large language models (LLMs) are increasingly used in decision support, it is important to understand whether their choices under uncertainty exhibit stable and interpretable be…
SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
Xuankun Rong, Wenke Huang, Tingfeng Wang +3
Multimodal large language models (MLLMs) have demonstrated impressive reasoning and instruction-following capabilities, yet their expanded modality space introduces new composition…
MultiVerse: A Multi-Turn Conversation Benchmark for Evaluating Large Vision and Language Models
Young-Jun Lee, Byung-Kwan Lee, Jianshu Zhang +9
Vision-and-Language Models (VLMs) have shown impressive capabilities on single-turn benchmarks, yet real-world applications often demand more intricate multi-turn dialogues. Existi…
MAPO: Mixed Advantage Policy Optimization
Wenke Huang, Quan Zhang, Yiyang Fang +11
Recent advances in reinforcement learning for foundation models, such as Group Relative Policy Optimization (GRPO), have significantly improved the performance of foundation models…