3 citations · 4 across the 5 of their papers we have counts for
5 papers
The Attacker Moves Second: Stronger Adaptive Attacks Bypass Defenses Against Llm Jailbreaks and Prompt Injections
Milad Nasr, Nicholas Carlini, Chawin Sitawarin +11
How should we evaluate the robustness of language model defenses? Current defenses against jailbreaks and prompt injections (which aim to prevent an attacker from eliciting harmful…
Text-to-Image Models Leave Identifiable Signatures: Implications for Leaderboard Security
Ali Naseh, Anshuman Suri, Yuefeng Peng +3
Generative AI leaderboards are central to evaluating model capabilities, but remain vulnerable to manipulation. Among key adversarial objectives is rank manipulation, where an atta…
Reasoning Introduces New Poisoning Attacks Yet Makes Them More Complicated
Hanna Foerster, Ilia Shumailov, Yiren Zhao +4
Early research into data poisoning attacks against Large Language Models (LLMs) demonstrated the ease with which backdoors could be injected. More recent LLMs add step-by-step reas…
R1dacted: Investigating Local Censorship in DeepSeek's R1 Language Model
Ali Naseh, Harsh Chaudhari, Jaechul Roh +3
DeepSeek recently released R1, a high-performing large language model (LLM) optimized for reasoning tasks. Despite its efficient training pipeline, R1 achieves competitive performa…
L3Cube-MahaSocialNER: A Social Media based Marathi NER Dataset and BERT models
Harsh Chaudhari, Anuja Patil, Dhanashree Lavekar +2
This work introduces the L3Cube-MahaSocialNER dataset, the first and largest social media dataset specifically designed for Named Entity Recognition (NER) in the Marathi language.…