When Dialects Collide: How Socioeconomic Mixing Affects Language Use
arXiv:2307.10016 · doi:10.1140/epjds/s13688-025-00563-9
Abstract
The socioeconomic background of people and how they use standard forms of language are not independent, as demonstrated in various sociolinguistic studies. However, the extent to which these correlations may be influenced by the mixing of people from different socioeconomic classes remains relatively unexplored from a quantitative perspective. In this work we leverage geotagged tweets and transferable computational methods to map deviations from standard English on a large scale, in seven thousand administrative areas of England and Wales. We combine these data with high-resolution income maps to assign a proxy socioeconomic indicator to home-located users. Strikingly, across eight metropolitan areas we find a consistent pattern suggesting that the more different socioeconomic classes mix, the less interdependent the frequency of their departures from standard grammar and their income become. Further, we propose an agent-based model of linguistic variety adoption that sheds light on the mechanisms that produce the observations seen in the data.
References in corpus (13)
- From mobile phone data to the spatial structure of cities
- The Twitter of Babel: Mapping World Languages through Microblogging Platforms
- Diffusion of Lexical Change in Social Media
- Mapping the Americanization of English in Space and Time
- Computational Socioeconomics
- Crowdsourcing Dialect Characterization through Twitter
- Dialectometric analysis of language variation in Twitter
- Storywrangler: A massive exploratorium for sociolinguistic, cultural, socioeconomic, and political timelines using Twitter
- Semantic homophily in online communication: evidence from Twitter
- Socioeconomic Dependencies of Linguistic Patterns in Twitter: A Multivariate Analysis
- Spatial evolution of human dialects
- Socioeconomic biases in urban mixing patterns of US metropolitan areas
- American cultural regions mapped through the lexical analysis of social media