large vision-language models 1multilingual safety 1multimodal alignment 1neuron-level alignment 1parameter-efficient fine-tuning 1
From the 1 of 14 linked papers with an AI index.
Showing cs.AIShow all
2 papers · 1 filter
cs.AI2026
No Single Neuron of Failure: Distributed Safety Alignment Against White-Box Attacks
Simiao Xie, Chuancheng Shi, Shangze Li +5
With the rapid release of open-weight large foundation models, safety threats are shifting from black-box jailbreaks to neuron-level white-box attacks that directly identify and ma…
cs.AI2026
One Anchor for All: Unified Multilingual and Multimodal Safety Alignment for LVLMs
Enyi Shi, Fei Shen, Chuancheng Shi +4
The paper introduces a neuron‑level safety alignment method that identifies and updates a tiny set of shared safety neurons across languages and modalities, enabling large vision‑l…