1 paper · 1 filter
Md Rysul Kabir, Zoran Tiganj
Open-weight language models can be rendered unsafe through several parameter-level interventions, yet models with matched harmful compliance can exhibit fundamentally different fai…