2 citations · 3 across the 14 of their papers we have counts for
1 paper · 1 filter
Long P. Hoang, Hai V. Le, Shaoyang Xu +2
Large language models (LLMs) are rigorously aligned to refuse harmful requests, a process that inherently cultivates a latent capacity to evaluate and recognize unsafe content. In…