3 citations · 10 across the 17 of their papers we have counts for
5 papers · 1 filter
ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization
Hujian Zhu, Yihao Huang, Felix Juefei-Xu +5
Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently em…
When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models
Juan Gabriel Kostelec, Qinghai Guo
Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality…
Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization
Yihao Huang, Chong Wang, Xiaojun Jia +5
Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts. The previous methods achieve…
Purifying Large Language Models by Ensembling a Small Language Model
Tianlin Li, Qian Liu, Tianyu Pang +4
The emerging success of large language models (LLMs) heavily relies on collecting abundant training data from external (untrusted) sources. Despite substantial efforts devoted to d…
Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One
Tianlin Li, Xiaoyu Zhang, Chao Du +5
The widespread adoption of large language models (LLMs) underscores the urgent need to ensure their fairness. However, LLMs frequently present dominant viewpoints while ignoring al…