1 paper · 1 filter
Zhaoxin Zhang, Borui Chen, Yiming Hu +3
Recent research on large language model (LLM) jailbreaks has primarily focused on techniques that bypass safety mechanisms to elicit overtly harmful outputs. However, such efforts…