1 paper
Hongyi Li, Jiawei Ye, Jie Wu +3
Large Language Models (LLMs) aligned with human feedback have recently garnered significant attention. However, it remains vulnerable to jailbreak attacks, where adversaries manipu…