1 paper · 1 filter
Tianyu Wu, Lingrui Mei, Ruibin Yuan +3
While recent advancements in large language model (LLM) alignment have enabled the effective identification of malicious objectives involving scene nesting and keyword rewriting, o…