Showing cs.CRShow all
2 papers · 1 filter
cs.CR2025
Jailbreaking Attacks vs. Content Safety Filters: How Far Are We in the LLM Safety Arms Race?
Yuan Xin, Dingfan Chen, Linyi Yang +2
As large language models (LLMs) are increasingly deployed, ensuring their safe use is paramount. Jailbreaking, adversarial prompts that bypass model alignment to trigger harmful ou…
cs.CR2024
Invisibility Cloak: Disappearance under Human Pose Estimation via Backdoor Attacks
Minxing Zhang, Wenshu Fan, Wenbo Jiang +3
Despite being significant in autonomous systems, Human Pose Estimation (HPE)'s potential risks to adversarial attacks have not received comparable attention with image classificati…