Prompt Perturbation in Retrieval-Augmented Generation based Large Language Models
arXiv:2402.07179 · doi:10.1145/3637528.3671932
Abstract
The robustness of large language models (LLMs) becomes increasingly important as their use rapidly grows in a wide range of domains. Retrieval-Augmented Generation (RAG) is considered as a means to improve the trustworthiness of text generation from LLMs. However, how the outputs from RAG-based LLMs are affected by slightly different inputs is not well studied. In this work, we find that the insertion of even a short prefix to the prompt leads to the generation of outputs far away from factually correct answers. We systematically evaluate the effect of such prefixes on RAG by introducing a novel optimization technique called Gradient Guided Prompt Perturbation (GGPP). GGPP achieves a high success rate in steering outputs of RAG-based LLMs to targeted wrong answers. It can also cope with instructions in the prompts requesting to ignore irrelevant context. We also exploit LLMs' neuron activation difference between prompts with and without GGPP perturbations to give a method that improves the robustness of RAG-based LLMs through a highly effective detector trained on neuron activation triggered by GGPP generated prompts. Our evaluation on open-sourced LLMs demonstrates the effectiveness of our methods.
12 pages, 9 figures
References in corpus (19)
- Explaining and Harnessing Adversarial Examples
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Improving language models by retrieving from trillions of tokens
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs
- Locating and Editing Factual Associations in GPT
- Large Language Models Can Be Easily Distracted by Irrelevant Context
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and Entailment
- On the Robustness of ChatGPT: An Adversarial and Out-of-distribution Perspective
- RAG vs Fine-tuning: Pipelines, Tradeoffs, and a Case Study on Agriculture
- PromptRobust: Towards Evaluating the Robustness of Large Language Models on Adversarial Prompts
- Mind the Gap: Assessing Temporal Generalization in Neural Language Models
- Re3: Generating Longer Stories With Recursive Reprompting and Revision
- NeMo Guardrails: A Toolkit for Controllable and Safe LLM Applications with Programmable Rails
- Dissecting Recall of Factual Associations in Auto-Regressive Language Models
- Vector Search with OpenAI Embeddings: Lucene Is All You Need
- Attention Satisfies: A Constraint-Satisfaction Lens on Factual Errors of Language Models
- How Does Generative Retrieval Scale to Millions of Passages?