From the 1 of 10 linked papers with an AI index.
1 paper · 1 filter
Cheng Wang, Zeming Wei, Qin Liu +1
Large Language Models (LLMs) can comply with harmful instructions, raising serious safety concerns despite their impressive capabilities. Recent work has leveraged probing-based ap…