1 paper
Xiaobing Sun, Perry Lam, Shaohua Li +4
Modern LLMs employ safety mechanisms that extend beyond surface-level input filtering to latent semantic representations and generation-time reasoning, enabling them to recover obf…