22 citations · 24 across the 3 of their papers we have counts for
7 papers
User Factor Adaptation for User Embedding via Multitask Learning
Xiaolei Huang, Michael J. Paul, Robin Burke +2
Language varies across users and their interested fields in social media data: words authored by a user across his/her interests may have different meanings (e.g., cool) or sentime…
Why Overfitting Isn't Always Bad: Retrofitting Cross-Lingual Word Embeddings to Dictionaries
Mozhi Zhang, Yoshinari Fujinuma, Michael J. Paul +1
Cross-lingual word embeddings (CLWE) are often evaluated on bilingual lexicon induction (BLI). Recent CLWE methods use linear projections, which underfit the training dictionary, t…
Multilingual Twitter Corpus and Baselines for Evaluating Demographic Bias in Hate Speech Recognition
Xiaolei Huang, Linzi Xing, Franck Dernoncourt +1
Existing research on fairness evaluation of document classification models mainly uses synthetic monolingual data without ground truth for author demographic attributes. In this wo…
Evaluating Topic Quality with Posterior Variability
Linzi Xing, Michael J. Paul, Giuseppe Carenini
Probabilistic topic models such as latent Dirichlet allocation (LDA) are popularly used with Bayesian inference methods such as Gibbs sampling to learn posterior distributions over…
An Empirical Study on Crosslingual Transfer in Probabilistic Topic Models
Shudong Hao, Michael J. Paul
Probabilistic topic modeling is a popular choice as the first step of crosslingual tasks to enable knowledge transfer and extract multilingual features. While many multilingual top…
Learning Multilingual Topics from Incomparable Corpus
Shudong Hao, Michael J. Paul
Multilingual topic models enable crosslingual tasks by extracting consistent topics from multilingual corpora. Most models require parallel or comparable training corpora, which li…