adversarial detection 1embedding space anisotropy 1geometric analysis 1multimodal robustness 1vision-language models 1
From the 1 of 9 linked papers with an AI index.
Showing cs.CLShow all
2 papers · 1 filter
cs.CL2025
Geometry-Guided Adversarial Prompt Detection via Curvature and Local Intrinsic Dimension
Canaan Yung, Hanxun Huang, Christopher Leckie +1
Adversarial prompts are capable of jailbreaking frontier large language models (LLMs) and inducing undesirable behaviours, posing a significant obstacle to their safe deployment. C…
cs.CL2025
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
Canaan Yung, Hadi Mohaghegh Dolatabadi, Sarah Erfani +1
Large language models (LLMs) are susceptible to social-engineered attacks that are human-interpretable but require a high level of comprehension for LLMs to counteract. Existing de…