The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
arXiv:2607.19292 · doi:10.1007/s43681-026-01132-0
Abstract
Current AI safety discourse still focuses disproportionately on visible failures, including obvious harms, dramatic misuse, and hypothetical catastrophic scenarios. That focus is incomplete. In deployed systems, many of the most consequential failures are quieter: plausible rather than spectacular, distributed across components rather than localized in a single output, and normalized by workflows before they are recognized as hazards. We argue that a central safety challenge in modern AI systems is increasingly not only whether a model emits a harmful response, but whether the broader socio-technical system preserves the conditions under which errors remain visible, contestable, containable, and recoverable. We propose a five-layer framework for diagnosing these hidden risks: (1) epistemic integrity, concerning whether evidence and uncertainty are represented honestly enough to support calibrated reliance; (2) control integrity, concerning whether authority, permissions, and action boundaries remain robust under attack and optimization; (3) temporal integrity, concerning whether safety holds across sessions, memory updates, and deployment drift; (4) organizational integrity, concerning whether institutions retain the capacity to audit, assign responsibility, and intervene effectively; and (5) ecosystem integrity, concerning whether AI systems preserve rather than erode the information environment on which future oversight depends. Across these layers, we identify under-recognized risk patterns, including overreliance, uncertainty and legitimacy laundering in retrieval, prompt injection, reward hacking, memory poisoning, evaluation deception, fictional human oversight, synthetic evidence pollution, and model collapse. We conclude with design and governance recommendations and a research agenda for shifting AI safety from model-centric evaluation toward socio-technical reliability.
Email by the arXiv Support Team: "Dear Gjergji, Thank you for your patience. Your appeal was accepted. You are welcome to resubmit this work at your convenience to cs.CY (cs.AI, cs.HC) Regards, arXiv Support"
References in corpus (27)
- To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making
- Towards Human-centered Explainable AI: A Survey of User Studies for Model Explanations
- Understanding accountability in algorithmic supply chains
- Black-Box Access is Insufficient for Rigorous AI Audits
- Fostering Appropriate Reliance on Large Language Models: The Role of Explanations, Sources, and Inconsistencies
- AI Safety for Everyone
- Alignment faking in large language models
- Trustworthiness in Retrieval-Augmented Generation Systems: A Survey
- Understanding the Effects of Miscalibrated AI Confidence on User Trust, Reliance, and Decision Efficacy
- Towards Safer Generative Language Models: A Survey on Safety Risks, Evaluations, and Improvements
- Large Language Models can Strategically Deceive their Users when Put Under Pressure
- Automatic and Universal Prompt Injection Attacks against Large Language Models
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Agentic Misalignment: How LLMs Could Be Insider Threats
- An Early Categorization of Prompt Injection Attacks on Large Language Models
- Defeating Prompt Injections by Design
- LLM Defenses Are Not Robust to Multi-Turn Human Jailbreaks Yet
- The BIG Argument for AI Safety Cases
- Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
- Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles
- A Closer Look at the Existing Risks of Generative AI: Mapping the Who, What, and How of Real-World Incidents
- Measuring and mitigating overreliance to build human-compatible AI
- How Memory Management Impacts LLM Agents: An Empirical Study of Experience-Following Behavior
- Unsafer in Many Turns: Benchmarking and Defending Multi-Turn Safety Risks in Tool-Using Agents
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Memory Poisoning Attack and Defense on Memory Based LLM-Agents
- Establishing Best Practices for Building Rigorous Agentic Benchmarks