2 papers
cs.CL2025
Universal Adversarial Suffixes for Language Models Using Reinforcement Learning with Calibrated Reward
Sampriti Soor, Suklav Ghosh, Arijit Sur
Language models are vulnerable to short adversarial suffixes that can reliably alter predictions. Previous works usually find such suffixes with gradient search or rule-based metho…
cs.CL2025
Universal Adversarial Suffixes Using Calibrated Gumbel-Softmax Relaxation
Sampriti Soor, Suklav Ghosh, Arijit Sur
Language models (LMs) are often used as zero-shot or few-shot classifiers by scoring label words, but they remain fragile to adversarial prompts. Prior work typically optimizes tas…