3 papers
cs.LG2025
Jailbreak-as-a-Service++: Unveiling Distributed AI-Driven Malicious Information Campaigns Powered by LLM Crowdsourcing
Yu Yan, Sheng Sun, Mingfeng Li +6
To prevent the misuse of Large Language Models (LLMs) for malicious purposes, numerous efforts have been made to develop the safety alignment mechanisms of LLMs. However, as multip…
cs.CL2025
Collaborative Stance Detection via Small-Large Language Model Consistency Verification
Yu Yan, Sheng Sun, Zixiang Tang +2
Stance detection on social media aims to identify attitudes expressed in tweets towards specific targets. Current studies prioritize Large Language Models (LLMs) over Small Languag…
cs.CL2024
Na'vi or Knave: Jailbreaking Language Models via Metaphorical Avatars
Yu Yan, Sheng Sun, Junqi Tong +2
Metaphor serves as an implicit approach to convey information, while enabling the generalized comprehension of complex subjects. However, metaphor can potentially be exploited to b…