artificial intelligence

Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities

arXiv:2607.13596

summary

The paper identifies and studies Protective Capacity Hallucination, where large language models claim to perform real-world protective actions they cannot actually do, especially in multi‑party dialogues, and shows this stems from gaps between role assignment and explicit capability boundaries.

Abstract

When cast as the protector of a vulnerable user yet given no explicit capability boundary, a large language model (LLM) may respond not by acknowledging its limits but by claiming to have taken -- or to be taking -- a real-world protective action it cannot perform, such as contacting emergency services or administering care. We term this phenomenon Protective Capacity Hallucination (PCH): a self-referential misattribution in which a model, acting in a protective role, asserts physical or institutional agency exceeding its affordances as a language model. In a three-phase study spanning eight LLMs and 13{,}600 sessions, we find PCH jointly gated by situational severity and interactional format: multi-party dialogic input drives it toward ceiling in most models across ordinary service domains, whereas in intimate-partner conflict -- a domain explicitly covered by safety alignment -- it remains at floor in all eight models despite greater physical severity. We interpret PCH as the signature of a deployment-design gap between role assignment and capability-boundary specification: a by-product of partial alignment in which a universally trained pressure to help outruns a domain-selective specification of how to help. Because suppression tracks alignment coverage rather than severity, deployment-side specification of capability boundaries emerges as a general mitigation target.

Topics & keywords

#protective capacity hallucination#llm alignment#hallucination#capability boundaries#safety#multi‑party dialoguelarge language modelprotective rolehallucinationalignmentcapability specificationempirical study
Protective Capacity Hallucination: When Large Language Models Claim Nonexistent Capabilities · wovepaper