2 papers
cs.CR2026
A Simple and Efficient Jailbreak Method Exploiting LLMs' Helpfulness
Xuan Luo, Yue Wang, Zefeng He +3
This study reveals a critical safety blind spot in modern LLMs: learning-style queries, which closely resemble ordinary educational questions, can reliably elicit harmful responses…
cs.LG2025
Multi-Task Model Merging via Adaptive Weight Disentanglement
Feng Xiong, Runxi Cheng, Wang Chen +4
Model merging has recently gained attention as an economical and scalable approach to incorporate task-specific weights from various tasks into a unified multi-task model. For exam…