activity
20232026
most citedYour Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One

3 citations · 10 across the 17 of their papers we have counts for

collaborators
Showing cs.CLShow all

5 papers · 1 filter

cs.CL2026

ICO: Enhancing Semantic-Shift Jailbreaks via Iterative Context Optimization

Hujian Zhu, Yihao Huang, Felix Juefei-Xu +5

Foundation models have achieved remarkable success across diverse tasks, but they remain vulnerable. To investigate such vulnerabilities, semantic-shift jailbreaks have recently em…

cs.CL2026

When Perplexity Lies: Generation-Focused Distillation of Hybrid Sequence Models

Juan Gabriel Kostelec, Qinghai Guo

Converting a pretrained Transformer into a more efficient hybrid model through distillation offers a promising approach to reducing inference costs. However, achieving high-quality…

cs.CL20241 cited

Efficient Universal Goal Hijacking with Semantics-guided Prompt Organization

Yihao Huang, Chong Wang, Xiaojun Jia +5

Universal goal hijacking is a kind of prompt injection attack that forces LLMs to return a target malicious response for arbitrary normal user prompts. The previous methods achieve…

cs.CL20242 cited

Purifying Large Language Models by Ensembling a Small Language Model

Tianlin Li, Qian Liu, Tianyu Pang +4

The emerging success of large language models (LLMs) heavily relies on collecting abundant training data from external (untrusted) sources. Despite substantial efforts devoted to d…

cs.CL20243 cited

Your Large Language Model is Secretly a Fairness Proponent and You Should Prompt it Like One

Tianlin Li, Xiaoyu Zhang, Chao Du +5

The widespread adoption of large language models (LLMs) underscores the urgent need to ensure their fairness. However, LLMs frequently present dominant viewpoints while ignoring al…