Where in the World are You? Geolocation and Language Identification in Twitter
arXiv:1308.0683 · doi:10.1080/00330124.2014.907699
Abstract
The movements of ideas and content between locations and languages are unquestionably crucial concerns to researchers of the information age, and Twitter has emerged as a central, global platform on which hundreds of millions of people share knowledge and information. A variety of research has attempted to harvest locational and linguistic metadata from tweets in order to understand important questions related to the 300 million tweets that flow through the platform each day. However, much of this work is carried out with only limited understandings of how best to work with the spatial and linguistic contexts in which the information was produced. Furthermore, standard, well-accepted practices have yet to emerge. As such, this paper studies the reliability of key methods used to determine language and location of content in Twitter. It compares three automated language identification packages to Twitter's user interface language setting and to a human coding of languages in order to identify common sources of disagreement. The paper also demonstrates that in many cases user-entered profile locations differ from the physical locations users are actually tweeting from. As such, these open-ended, user-generated, profile locations cannot be used as useful proxies for the physical locations from which information is published to Twitter.
Cited by in corpus (18)
- The Effects of Twitter Sentiment on Stock Price Returns
- The Twitter of Babel: Mapping World Languages through Microblogging Platforms
- A Survey of Location Prediction on Twitter
- Estimating Local Commuting Patterns From Geolocated Twitter Data
- A Biased Review of Biases in Twitter Studies on Political Collective Action
- Socioeconomic Dependencies of Linguistic Patterns in Twitter: A Multivariate Analysis
- Global Syntactic Variation in Seven Languages: Towards a Computational Dialectology
- Cross-language sentiment analysis of European Twitter messages duringthe COVID-19 pandemic
- Mapping Languages: The Corpus of Global Language Use
- Choosing the right home location definition method for the given dataset
- Mega-COV: A Billion-Scale Dataset of 100+ Languages for COVID-19
- A Hierarchical Location Prediction Neural Network for Twitter User Geolocation
- Leave no Place Behind: Improved Geolocation in Humanitarian Documents
- On-the-fly Detection of Autogenerated Tweets
- Gender Imbalance and Spatiotemporal Patterns of Contributions to Citizen Science Projects: the case of Zooniverse
- Infringement of Tweets Geo-Location Privacy: an approach based on Graph Convolutional Neural Networks
- Revealing the Global Linguistic and Geographical Disparities of Public Awareness to Covid-19 Outbreak through Social Media
- Novel Keyword Extraction and Language Detection Approaches