4 papers · 1 filter
from Benign import Toxic: Jailbreaking the Language Model via Adversarial Metaphors
Yu Yan, Sheng Sun, Zenghao Duan +5
Current studies have exposed the risk of Large Language Models (LLMs) generating harmful content by jailbreak attacks. However, they overlook that the direct generation of harmful…
SearchAttack: Red-Teaming LLMs against Knowledge-to-Action Threats under Online Web Search
Yu Yan, Sheng Sun, Mingfeng Li +6
Recently, people have suffered from LLM hallucination and have become increasingly aware of the reliability gap of LLMs in open and knowledge-intensive tasks. As a result, they hav…
Collaborative Stance Detection via Small-Large Language Model Consistency Verification
Yu Yan, Sheng Sun, Zixiang Tang +2
Stance detection on social media aims to identify attitudes expressed in tweets towards specific targets. Current studies prioritize Large Language Models (LLMs) over Small Languag…
Na'vi or Knave: Jailbreaking Language Models via Metaphorical Avatars
Yu Yan, Sheng Sun, Junqi Tong +2
Metaphor serves as an implicit approach to convey information, while enabling the generalized comprehension of complex subjects. However, metaphor can potentially be exploited to b…