Corporate Loyalty: Some AI Systems Differentially Downplay their Creators' Controversies
arXiv:2609.00373
Abstract
Language models have become a major mediator of politically relevant information and are used to assist decision-making in high-stakes settings. Due to their wide use, the developers of popular AI systems have a powerful ability to subtly influence the marketplace of ideas. Recognizing this, many AI companies have publicly discussed the importance of AI systems not taking positions or disseminating information in ways that favor special interests. In this paper, we ask whether popular AI systems have a tendency to downplay the controversies associated with the companies that created them. In a pre-registered experiment, we elicit open-ended discussions from 21 models from 7 companies on 206 negative news stories using 25 prompt templates to assess how favorably each model discusses controversies from each company. We find strong evidence (p<10^-5) that models from xAI, DeepSeek, Anthropic, and OpenAI tend to discuss controversies from their respective companies in a differentially positive way compared to others. We find no such evidence for Alibaba, Meta, and Google. Finally, we conclude with a discussion of the differing implications of whether these behaviors were intentionally given to models by developers, unintentionally given to models by developers, or represent a form of emergent misalignment.
Preregistered on OSF: https://osf.io/twqps/overview