Overcoming Language Variation in Sentiment Analysis with Social Attention
arXiv:1511.06052
Abstract
Variation in language is ubiquitous, particularly in newer forms of writing such as social media. Fortunately, variation is not random, it is often linked to social properties of the author. In this paper, we show how to exploit social networks to make sentiment analysis more robust to social language variation. The key idea is linguistic homophily: the tendency of socially linked individuals to use language in similar ways. We formalize this idea in a novel attention-based neural network architecture, in which attention is divided among several basis models, depending on the author's position in the social network. This has the effect of smoothing the classification function across the social network, and makes it possible to induce personalized classifiers even for authors for whom there is no labeled data or demographic metadata. This model significantly improves the accuracies of sentiment analysis on Twitter and on review data.
Published in Transactions of the Association for Computational Linguistics (TACL), 2017. Please cite the TACL version: https://transacl.org/ojs/index.php/tacl/article/view/1024
References in corpus (5)
Cited by in corpus (7)
- Characterization of Time-variant and Time-invariant Assessment of Suicidality on Reddit using C-SSRS
- Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora
- Analysis of Twitter Users' Lifestyle Choices using Joint Embedding Model
- Twitter User Representation Using Weakly Supervised Graph Embedding
- Toward Socially-Infused Information Extraction: Embedding Authors, Mentions, and Entities
- Generalisation in Named Entity Recognition: A Quantitative Analysis
- Expanding Subjective Lexicons for Social Media Mining with Embedding Subspaces