8 papers
AUDETER: A Large-scale Dataset for Deepfake Audio Detection in Open Worlds
Qizhou Wang, Hanxun Huang, Guansong Pang +2
Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, th…
Toward Universal and Transferable Jailbreak Attacks on Vision-Language Models
Kaiyuan Cui, Yige Li, Yutao Wu +4
Vision-language models (VLMs) extend large language models (LLMs) with vision encoders, enabling text generation conditioned on both images and text. However, this multimodal integ…
Geometry-Guided Adversarial Prompt Detection via Curvature and Local Intrinsic Dimension
Canaan Yung, Hanxun Huang, Christopher Leckie +1
Adversarial prompts are capable of jailbreaking frontier large language models (LLMs) and inducing undesirable behaviours, posing a significant obstacle to their safe deployment. C…
Multi-level Certified Defense Against Poisoning Attacks in Offline Reinforcement Learning
Shijie Liu, Andrew C. Cullen, Paul Montague +2
Similar to other machine learning frameworks, Offline Reinforcement Learning (RL) is shown to be vulnerable to poisoning attacks, due to its reliance on externally sourced datasets…
Round Trip Translation Defence against Large Language Model Jailbreaking Attacks
Canaan Yung, Hadi Mohaghegh Dolatabadi, Sarah Erfani +1
Large language models (LLMs) are susceptible to social-engineered attacks that are human-interpretable but require a high level of comprehension for LLMs to counteract. Existing de…
Intrinsic and Extrinsic Factor Disentanglement for Recommendation in Various Context Scenarios
Yixin Su, Wei Jiang, Fangquan Lin +6
In recommender systems, the patterns of user behaviors (e.g., purchase, click) may vary greatly in different contexts (e.g., time and location). This is because user behavior is jo…