7 citations · 13 across the 15 of their papers we have counts for
Showing cs.CRShow all
2 papers · 1 filter
cs.CR2024
Global Challenge for Safe and Secure LLMs Track 1
Xiaojun Jia, Yihao Huang, Yang Liu +27
This paper introduces the Global Challenge for Safe and Secure Large Language Models (LLMs), a pioneering initiative organized by AI Singapore (AISG) and the CyberSG R&D Programme…
cs.CR2024★ 2 cited
From Theft to Bomb-Making: The Ripple Effect of Unlearning in Defending Against Jailbreak Attacks
Zhexin Zhang, Junxiao Yang, Yida Lu +5
Large Language Models (LLMs) are known to be vulnerable to jailbreak attacks. An important observation is that, while different types of jailbreak attacks can generate significantl…