8 citations · 8 across the 1 of their papers we have counts for
4 papers
E-M3RF: An Equivariant Multimodal 3D Re-assembly Framework
Adeela Islam, Stefano Fiorini, Manuel Lecha +4
3D reassembly is a fundamental geometric problem, and in recent years it has increasingly been challenged by deep learning methods rather than classical optimization. While learnin…
ALMANACS: A Simulatability Benchmark for Language Model Explainability
Edmund Mills, Shiye Su, Stuart Russell +1
How do we measure the efficacy of language model explainability methods? While many explainability methods have been developed, they are typically evaluated on bespoke tasks, preve…
Tensor Trust: Interpretable Prompt Injection Attacks from an Online Game
Sam Toyer, Olivia Watkins, Ethan Adrian Mendes +9
While Large Language Models (LLMs) are increasingly being used in real-world applications, they remain vulnerable to prompt injection attacks: malicious third party prompts that su…
Image Hijacks: Adversarial Images can Control Generative Models at Runtime
Luke Bailey, Euan Ong, Stuart Russell +1
Are foundation models secure against malicious actors? In this work, we focus on the image input to a vision-language model (VLM). We discover image hijacks, adversarial images tha…