16 citations · 16 across the 1 of their papers we have counts for
3 papers
DICT-MLM: Improved Multilingual Pre-Training using Bilingual Dictionaries
Aditi Chaudhary, Karthik Raman, Krishna Srinivasan +1
Pre-trained multilingual language models such as mBERT have shown immense gains for several natural language processing (NLP) tasks, especially in the zero-shot cross-lingual setti…
Google COVID-19 Search Trends Symptoms Dataset: Anonymization Process Description (version 1.0)
Shailesh Bavadekar, Andrew Dai, John Davis +27
This report describes the aggregation and anonymization process applied to the initial version of COVID-19 Search Trends symptoms dataset (published at https://goo.gle/covid19sympt…
Learning Multilingual Word Embeddings Using Image-Text Data
Karan Singhal, Karthik Raman, Balder ten Cate
There has been significant interest recently in learning multilingual word embeddings -- in which semantically similar words across languages have similar embeddings. State-of-the-…