3 citations · 3 across the 3 of their papers we have counts for
3 papers
cs.LG2025★ 3 cited
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Milad Nasr, Nicholas Carlini, Chawin Sitawarin +11
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful…
cs.CR2025
Lessons from Defending Gemini Against Indirect Prompt Injections
Chongyang Shi, Sharon Lin, Shuang Song +11
Gemini is increasingly used to perform tasks on behalf of users, where function-calling and tool-use capabilities enable the model to access user data. Some tools, however, require…
cs.LG2025
Large Language Models Can Verbatim Reproduce Long Malicious Sequences
Sharon Lin, Krishnamurthy, Dvijotham +4
Backdoor attacks on machine learning models have been extensively studied, primarily within the computer vision domain. Originally, these attacks manipulated classifiers to generat…