1 paper
Long P. Hoang, Hai V. Le, Shaoyang Xu +2
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In…