Showing cs.CRShow all
3 papers · 1 filter
cs.CR2025
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Yuan Xin, Dingfan Chen, Linyi Yang +2
As large language models (LLMs) are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful ou…
cs.CR2025
Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks
Minxing Zhang, Wenshu Fan, Wenbo Jiang +3
Despite being significant in autonomous systems, Human Pose Estimation (HPE)'s potential risks to adversarial attacks have not received comparable attention with image classificati…
cs.CR2024
Transferable Availability Poisoning Attacks
Yiyong Liu, Michael Backes, Xiao Zhang
We consider availability data poisoning attacks, where an adversary aims to degrade the overall test accuracy of a machine learning model by crafting small perturbations to its tra…