1 paper
Ramakrishna P. Kompella, Aadit Mahajan
Safety alignment in multilingual models is uneven: a model that reliably refuses a harmful request in English will often comply with the same request in a lower-resource language.…