1 paper · 1 filter
Md Mokarram Chowdhury, Ernie Chang, Yang Li
Large language models are trained to follow instructions while refusing harmful requests. Jailbreaks exploit this balance to elicit content a model would ordinarily reject. Rolepla…