4 papers · 1 filter
"I am bad": Interpreting Stealthy, Universal and Robust Audio Jailbreaks in Audio-Language Models
Isha Gupta, David Khachaturov, Robert Mullins
The rise of multimodal large language models has introduced innovative human-machine interaction paradigms but also significant challenges in machine learning safety. Audio-Languag…
Adversarial Suffix Filtering: a Defense Pipeline for LLMs
David Khachaturov, Robert Mullins
Large Language Models (LLMs) are increasingly embedded in autonomous systems and public-facing environments, yet they remain susceptible to jailbreak vulnerabilities that may under…
Watermarking Needs Input Repetition Masking
David Khachaturov, Robert Mullins, Ilia Shumailov +1
Recent advancements in Large Language Models (LLMs) raised concerns over potential misuse, such as for spreading misinformation. In response two counter measures emerged: machine l…
Complexity Matters: Effective Dimensionality as a Measure for Adversarial Robustness
David Khachaturov, Robert Mullins
Quantifying robustness in a single measure for the purposes of model selection, development of adversarial training methods, and anticipating trends has so far been elusive. The si…