natural language processing

STEREODISCO: Discovering Stereotypicality in LLMs

arXiv:2607.27824

summary

The paper introduces STEREODISCO, a framework that discovers stereotypical semantic axes in large language models by probing their activation spaces using WordNet antonym pairs, and demonstrates its use on LLAMA-3 and MISTRAL models, uncovering novel stereotype dimensions.

Abstract

LLMs encode, convey, and perpetuate stereotypes. Prior computational research focuses on a small set of semantic axes investigated in social psychology, and operates on word embeddings produced by language models, leaving open which other semantic axes carry stereotypical associations in LLMs and how LLMs internally represent such axes. We introduce STEREODISCO, a framework that adapts the semantic differential method (Osgood et al., 1957) to the systematic study of stereotypes in LLM internal representations. STEREODISCO constructs approx. 2,000 candidate semantic axes from WordNet antonym synsets, recovers each as a geometric axis in the LLM's activation space via probing, and identifies stereotypical axes via a statistical test over concept projections. As a case study, we apply STEREODISCO to social group stereotypes with LLAMA-3-8B-INSTRUCT and MISTRAL-7B-INSTRUCT. We find that the two LLMs agree with each other on social group ratings more than with humans, suggesting that LLM-encoded stereotype content diverges from that documented in social psychology. We also discover stereotypical axes not investigated in prior work -- including humble vs. proud, narrow-minded vs. broad-minded, and cowardly vs. brave, which human annotators independently confirm.

Topics & keywords

#stereotype detection#large language models#semantic axes#probing#bias analysissemantic differentialWordNet antonym synsetsactivation space probingLLAMA-3-8B-INSTRUCTMISTRAL-7B-INSTRUCT
STEREODISCO: Discovering Stereotypicality in LLMs · wovepaper