2 papers
cs.AI2026
When Refusal Looks Safe: The Refusal-Cue Shortcut in Safety Guard Models
Yu Feng, Chunting Zang, Chen Shen +4
Safety guards are widely used to filter harmful content and are typically trained via supervised fine-tuning on labeled prompt-response pairs. We audit two widely used safety-guard…
cs.CL2025
Improving Complex Reasoning with Dynamic Prompt Corruption: A soft prompt Optimization Approach
Sinan Fan, Liang Xie, Chen Shen +7
Prompt-tuning (PT) for large language models (LLMs) can facilitate the performance on various conventional NLP tasks with significantly fewer trainable parameters. However, our inv…