Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
arXiv:2505.21091 · doi:10.1145/3715275.3732038
Abstract
System prompts in Large Language Models (LLMs) are predefined directives that guide model behaviour, taking precedence over user inputs in text processing and generation. LLM deployers increasingly use them to ensure consistent responses across contexts. While model providers set a foundation of system prompts, deployers and third-party developers can append additional prompts without visibility into others' additions, while this layered implementation remains entirely hidden from end-users. As system prompts become more complex, they can directly or indirectly introduce unaccounted for side effects. This lack of transparency raises fundamental questions about how the position of information in different directives shapes model outputs. As such, this work examines how the placement of information affects model behaviour. To this end, we compare how models process demographic information in system versus user prompts across six commercially available LLMs and 50 demographic groups. Our analysis reveals significant biases, manifesting in differences in user representation and decision-making scenarios. Since these variations stem from inaccessible and opaque system-level configurations, they risk representational, allocative and potential other biases and downstream harms beyond the user's ability to detect or correct. Our findings draw attention to these critical issues, which have the potential to perpetuate harms if left unexamined. Further, we argue that system prompt analysis must be incorporated into AI auditing processes, particularly as customisable system prompts become increasingly prevalent in commercial AI deployments.
Published in Proceedings of ACM FAccT 2025 Update Comment: Fixed the error where user vs. system and implicit vs. explicit labels in the heatmaps were switched. The takeaways remain the same
References in corpus (32)
- Training language models to follow instructions with human feedback
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- On the Opportunities and Risks of Foundation Models
- Human heuristics for AI-generated language are flawed
- BOLD: Dataset and Metrics for Measuring Biases in Open-Ended Language Generation
- Who Audits the Auditors? Recommendations from a field scan of the algorithmic auditing ecosystem
- Dislocated Accountabilities in the AI Supply Chain: Modularity and Developers' Notions of Responsibility
- Understanding accountability in algorithmic supply chains
- What's Sex Got To Do With Fair Machine Learning?
- Bias Against 93 Stigmatized Groups in Masked Language Models and Downstream Sentiment Classification Tasks
- Participation in the age of foundation models
- Large Language Models Portray Socially Subordinate Groups as More Homogeneous, Consistent with a Bias Observed in Humans
- "I'm fully who I am": Towards Centering Transgender and Non-Binary Voices to Measure Biases in Open Language Generation
- Out of Context: Investigating the Bias and Fairness Concerns of "Artificial Intelligence as a Service"
- A Framework for Fairness: A Systematic Review of Existing Fair AI Solutions
- Beyond Behaviorist Representational Harms: A Plan for Measurement and Mitigation
- On the Effectiveness of Adapter-based Tuning for Pretrained Language Model Adaptation
- When "A Helpful Assistant" Is Not Really Helpful: Personas in System Prompts Do Not Improve Performances of Large Language Models
- The Unreasonable Effectiveness of Eccentric Automatic Prompts
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMs
- On LLM Wizards: Identifying Large Language Models' Behaviors for Wizard of Oz Experiments
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor Dataset
- Measuring and Controlling Instruction (In)Stability in Language Model Dialogs
- Aligning to Thousands of Preferences via System Message Generalization
- Can LLMs Follow Simple Rules?
- See It from My Perspective: How Language Affects Cultural Bias in Image Understanding
- You don't need a personality test to know these models are unreliable: Assessing the Reliability of Large Language Models on Psychometric Instruments
- SysBench: Can Large Language Models Follow System Messages?
- Unveiling and Mitigating Bias in Large Language Model Recommendations: A Path to Fairness
- PromptKeeper: Safeguarding System Prompts for LLMs
- Mixture-of-Instructions: Aligning Large Language Models via Mixture Prompting
- RelayAttention for Efficient Large Language Model Serving with Long System Prompts