Showing cs.LGShow all
2 papers · 1 filter
cs.LG2025
Soft Prompt Threats: Attacking Safety Alignment and Unlearning in Open-Source LLMs through the Embedding Space
Leo Schwinn, David Dobre, Sophie Xhonneux +2
Current research in adversarial robustness of LLMs focuses on discrete input manipulations in the natural language space, which can be directly transferred to closed-source models.…
cs.LG2024
In-Context Learning Can Re-learn Forbidden Tasks
Sophie Xhonneux, David Dobre, Jian Tang +2
Despite significant investment into safety training, large language models (LLMs) deployed in the real world still suffer from numerous vulnerabilities. One perspective on LLM safe…